{
  "id": 243343,
  "title": "Part of the 18th place solution (Iafoss's model)",
  "url": "/competitions/birdclef-2021/writeups/just-do-it-part-of-the-18th-place-solution-iafoss-",
  "author_name": "",
  "post_date": "2021-06-10T03:16:39.950Z",
  "votes": 30,
  "comment_count": 17,
  "views": 0,
  "content": "<h2>Summary</h2>\n<ul>\n<li>Sequence based model trained on weak labels</li>\n<li>Multi-head self-attention applied to entire sequences</li>\n<li>Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.</li>\n<li>Rainforest noise</li>\n<li>postprocessing based on 2.5s shifted sequence</li>\n<li>best <strong>single model performance 0.7598/0.6486</strong> at public/private LB</li>\n</ul>\n<h2>Introduction</h2>\n<p>Congratulation to all participants. To begin with, our team would like to thank organizers and Kaggle team for making this competition possible. I also would like to express my gratitude to my teammates: <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> , <a href=\"https://www.kaggle.com/whurobin\" target=\"_blank\">@whurobin</a> , <a href=\"https://www.kaggle.com/philipkd\" target=\"_blank\">@philipkd</a> , and especially to <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> : he helped me today until 3:30 AM of his time to assemble the final submissions. Below I'll share the key point of our approach, and <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> may publish a separate post describing specific details of his part.<br>\nThis competition was relatively short for me:  I had just ~2 weeks after finishing another competition. Below I provide the key points of my portion of our solution.</p>\n<h2>Model</h2>\n<p>My part of the solution is largely influenced by the method <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183258\" target=\"_blank\">I proposed in 2020 Cornell Birdcall Identification challene</a> targeted at building a segmentation model based on weak classification labels.</p>\n<p><img src=\"https://i.ibb.co/1bZ0Ghg/image.png\" alt=\"\"></p>\n<p>I collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and then apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp (LSE) pooling to produce the prediction global prediction. In this year I quite modified this approach, and the model predictions look much better in comparison with ones from the previous year. First, following <a href=\"https://arxiv.org/pdf/1411.6228.pdf\" target=\"_blank\">this article</a> I introduced temperature to LSE equal to 5. Another trick is enforcing the minimum of predicted sequence to be 0: bird call is only present in a short time frame. It is achieved by an additional LSE output with -1 temperature followed by applying a loss with zero GT label. These modifications drastically boosted the model performance. Another important thing is using sufficiently short audio segments during training. Long segments help to overcome label weakness, but I realized that a model trained on long segments learns only easy sounds ignoring the rest (it is often just enough to find an easy example in a long sequence to assign the corresponding label). To encourage the model to learn difficult examples I used 5s chunks during training (but I used PL based sampling to deal with weak labels). Important component of the model is a multi-head self-attention applied to entire sequence:  in my experiments this component boosted F1 of sequence based predictions (10 min) by 0.01-0.02 in comparison with CV evaluated based solely on 5s chunks.</p>\n<p><strong>Additional details:</strong><br>\nBackbone: ResNeXt50<br>\nLoss: Focal loss, corrected to be suitable for soft labels<br>\nAugmentation: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183199\" target=\"_blank\">MixUp with max label</a> (proposed by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> ), pink noise, external noise<br>\n<strong>[important]</strong> n_ftt = 2048. It appeared that models trained with n_ftt = 1024 are very susceptible to the domain shift, and n_ftt = 2048 leads to much better performance on soundscapes.</p>\n<h2>Train-Segment-Shift</h2>\n<p>It is the concept that was used to address the competition problem. <strong>There are 2 major issues to address in this challenge: (1) label weakness and noise, and (2) domain mismatch</strong> between train and test data. My approach consists of 3 steps: train, segment, and shift.</p>\n<p>(1) I begin with creating a segmentation model based on the original labels assigned to entire clips. The main difficulty is that clip labels are not equal to chunk labels because the clips may have no bird calls in particular parts or bird calls from other species. So <strong>direct training of a model on clip labels assigned to chunks cannot produce a good model</strong>, but it gives PL for the next step.</p>\n<p>(2) Run similar setup as step 1 but do sampling based on PL segmentation, making sure that the primary label bird call is included in each selected audio segment. In addition, I applied segmentation loss to region having high confidence pseudo labels, while masking regions where confidence is not sufficient. It is done to guide the model, and it improved the quality of segmentation labels at the end. <br>\nThe image below shows the model predictions from the same audio for top 5 classes. Blue color corresponds to the primary label, while orange corresponds to the secondary label. The bottom image shows an example of segmentation PL used at step2 and step3. Red bars indicate high confidence PL, while orange ones correspond to low confidence PL (used for masking).<br>\n<img src=\"https://i.ibb.co/SXqcYw6/image.png\" alt=\"\"><br>\nAt this step I got ~0.85 F1 CV (based on short audio) evaluated by taking max values for predictions within entire clip train files (so the provided weak labels can be used for model evaluation). However, if I try to apply this model directly to train soundscapes (which are similar to test data), I get only 0.67 CV (soundscapes), which indicated the domain mismatch between short audio and soundscapes.</p>\n<p>(3) That's why the next step is crucial: domain shift accommodation. In this competition the mismatch is not as bad as one year ago (in that case soundscapes were recorded with 16 kHz rate only). Though, if you plot the spectrograms or listen soundscapes, they are still drastically different from short train clips. <br>\nInitially I tried to add noise extracted from train soundscapes (with proper CV split) to each loaded audio segment (I didn't use train soundscapes directly, only noise from them) + pink noise. It didn't work well giving 0.7231/0.6401 at public/private LB. Next, I realized that using <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/data\" target=\"_blank\">Rain Forest data</a> works really well to eliminate FP. This data contains multiple background sounds and sounds similar to bird calls but not bird calls. <strong>So the model learns to ignore sounds similar to bird calls as well as bird calls from unknown species.</strong> This approach boosted the single model performance to  0.7598/0.6486 at public/private LB.</p>\n<p>For example, the image below shows a single model single fold predictions for the first two minutes of 10534_SSW_20170429.ogg soundscape for top 10 classes. The predictions look really confident for detecting nocalls. There is some confusion between some classes though. Also, I noticed that GT labels missed several calls: it is quite clear that the calls are present, but probably experts were not confident enough about the species because of the noise.</p>\n<p><img src=\"https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png\" alt=\"\"><br>\nOne may ask, why do we need step2 and high accuracy segmentation PL. Without noise the model is able to localize the areas with corresponding bird calls. Meanwhile if extensive noise is added, the model needs hints on localization of birdcalls and ignoring the rest. Segmentation PL help to guide training.</p>\n<h2>Postprocessing</h2>\n<p>One important portion of our solution is postprocessing. We averaged the prediction for a given 5s chunk with predictions for chunks shifted by 2.5 seconds forward and backward (one can consider a shifted audio), so <code>p = 0.5*p0 + 0.25*p_r + 0.25*pl</code>. This postprocessing is especially effective to deal with cases when a birdcall is located at a chunk boundary. In addition, <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> was using a postprocessing based on multiplying all chunk predictions for a given class by 1.3 if the maximum predicted value for the given class over the entire sequence is reaching a particular threshold.</p>\n<h2>Ensembling</h2>\n<p>Our best final submission is combination of 2 models having 0.7471/0.6206 and 0.7598/0.6486 public LB and private LB scores which gives 0.7736/0.6605 for weighted ensemble.</p>",
  "messages": [
    {
      "id": "1332447",
      "postDate": "06/02/2021 06:08:33",
      "content": "<h2>Summary</h2>\n<ul>\n<li>Sequence based model trained on weak labels</li>\n<li>Multi-head self-attention applied to entire sequences</li>\n<li>Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.</li>\n<li>Rainforest noise</li>\n<li>postprocessing based on 2.5s shifted sequence</li>\n<li>best <strong>single model performance 0.7598/0.6486</strong> at public/private LB</li>\n</ul>\n<h2>Introduction</h2>\n<p>Congratulation to all participants. To begin with, our team would like to thank organizers and Kaggle team for making this competition possible. I also would like to express my gratitude to my teammates: <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> , <a href=\"https://www.kaggle.com/whurobin\" target=\"_blank\">@whurobin</a> , <a href=\"https://www.kaggle.com/philipkd\" target=\"_blank\">@philipkd</a> , and especially to <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> : he helped me today until 3:30 AM of his time to assemble the final submissions. Below I'll share the key point of our approach, and <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> may publish a separate post describing specific details of his part.<br>\nThis competition was relatively short for me:  I had just ~2 weeks after finishing another competition. Below I provide the key points of my portion of our solution.</p>\n<h2>Model</h2>\n<p>My part of the solution is largely influenced by the method <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183258\" target=\"_blank\">I proposed in 2020 Cornell Birdcall Identification challene</a> targeted at building a segmentation model based on weak classification labels.</p>\n<p><img src=\"https://i.ibb.co/1bZ0Ghg/image.png\" alt=\"\"></p>\n<p>I collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and then apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp (LSE) pooling to produce the prediction global prediction. In this year I quite modified this approach, and the model predictions look much better in comparison with ones from the previous year. First, following <a href=\"https://arxiv.org/pdf/1411.6228.pdf\" target=\"_blank\">this article</a> I introduced temperature to LSE equal to 5. Another trick is enforcing the minimum of predicted sequence to be 0: bird call is only present in a short time frame. It is achieved by an additional LSE output with -1 temperature followed by applying a loss with zero GT label. These modifications drastically boosted the model performance. Another important thing is using sufficiently short audio segments during training. Long segments help to overcome label weakness, but I realized that a model trained on long segments learns only easy sounds ignoring the rest (it is often just enough to find an easy example in a long sequence to assign the corresponding label). To encourage the model to learn difficult examples I used 5s chunks during training (but I used PL based sampling to deal with weak labels). Important component of the model is a multi-head self-attention applied to entire sequence:  in my experiments this component boosted F1 of sequence based predictions (10 min) by 0.01-0.02 in comparison with CV evaluated based solely on 5s chunks.</p>\n<p><strong>Additional details:</strong><br>\nBackbone: ResNeXt50<br>\nLoss: Focal loss, corrected to be suitable for soft labels<br>\nAugmentation: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183199\" target=\"_blank\">MixUp with max label</a> (proposed by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> ), pink noise, external noise<br>\n<strong>[important]</strong> n_ftt = 2048. It appeared that models trained with n_ftt = 1024 are very susceptible to the domain shift, and n_ftt = 2048 leads to much better performance on soundscapes.</p>\n<h2>Train-Segment-Shift</h2>\n<p>It is the concept that was used to address the competition problem. <strong>There are 2 major issues to address in this challenge: (1) label weakness and noise, and (2) domain mismatch</strong> between train and test data. My approach consists of 3 steps: train, segment, and shift.</p>\n<p>(1) I begin with creating a segmentation model based on the original labels assigned to entire clips. The main difficulty is that clip labels are not equal to chunk labels because the clips may have no bird calls in particular parts or bird calls from other species. So <strong>direct training of a model on clip labels assigned to chunks cannot produce a good model</strong>, but it gives PL for the next step.</p>\n<p>(2) Run similar setup as step 1 but do sampling based on PL segmentation, making sure that the primary label bird call is included in each selected audio segment. In addition, I applied segmentation loss to region having high confidence pseudo labels, while masking regions where confidence is not sufficient. It is done to guide the model, and it improved the quality of segmentation labels at the end. <br>\nThe image below shows the model predictions from the same audio for top 5 classes. Blue color corresponds to the primary label, while orange corresponds to the secondary label. The bottom image shows an example of segmentation PL used at step2 and step3. Red bars indicate high confidence PL, while orange ones correspond to low confidence PL (used for masking).<br>\n<img src=\"https://i.ibb.co/SXqcYw6/image.png\" alt=\"\"><br>\nAt this step I got ~0.85 F1 CV (based on short audio) evaluated by taking max values for predictions within entire clip train files (so the provided weak labels can be used for model evaluation). However, if I try to apply this model directly to train soundscapes (which are similar to test data), I get only 0.67 CV (soundscapes), which indicated the domain mismatch between short audio and soundscapes.</p>\n<p>(3) That's why the next step is crucial: domain shift accommodation. In this competition the mismatch is not as bad as one year ago (in that case soundscapes were recorded with 16 kHz rate only). Though, if you plot the spectrograms or listen soundscapes, they are still drastically different from short train clips. <br>\nInitially I tried to add noise extracted from train soundscapes (with proper CV split) to each loaded audio segment (I didn't use train soundscapes directly, only noise from them) + pink noise. It didn't work well giving 0.7231/0.6401 at public/private LB. Next, I realized that using <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/data\" target=\"_blank\">Rain Forest data</a> works really well to eliminate FP. This data contains multiple background sounds and sounds similar to bird calls but not bird calls. <strong>So the model learns to ignore sounds similar to bird calls as well as bird calls from unknown species.</strong> This approach boosted the single model performance to  0.7598/0.6486 at public/private LB.</p>\n<p>For example, the image below shows a single model single fold predictions for the first two minutes of 10534_SSW_20170429.ogg soundscape for top 10 classes. The predictions look really confident for detecting nocalls. There is some confusion between some classes though. Also, I noticed that GT labels missed several calls: it is quite clear that the calls are present, but probably experts were not confident enough about the species because of the noise.</p>\n<p><img src=\"https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png\" alt=\"\"><br>\nOne may ask, why do we need step2 and high accuracy segmentation PL. Without noise the model is able to localize the areas with corresponding bird calls. Meanwhile if extensive noise is added, the model needs hints on localization of birdcalls and ignoring the rest. Segmentation PL help to guide training.</p>\n<h2>Postprocessing</h2>\n<p>One important portion of our solution is postprocessing. We averaged the prediction for a given 5s chunk with predictions for chunks shifted by 2.5 seconds forward and backward (one can consider a shifted audio), so <code>p = 0.5*p0 + 0.25*p_r + 0.25*pl</code>. This postprocessing is especially effective to deal with cases when a birdcall is located at a chunk boundary. In addition, <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> was using a postprocessing based on multiplying all chunk predictions for a given class by 1.3 if the maximum predicted value for the given class over the entire sequence is reaching a particular threshold.</p>\n<h2>Ensembling</h2>\n<p>Our best final submission is combination of 2 models having 0.7471/0.6206 and 0.7598/0.6486 public LB and private LB scores which gives 0.7736/0.6605 for weighted ensemble.</p>",
      "rawMarkdown": "## Summary\n- Sequence based model trained on weak labels\n- Multi-head self-attention applied to entire sequences\n- Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.\n- Rainforest noise\n- postprocessing based on 2.5s shifted sequence\n- best **single model performance 0.7598/0.6486** at public/private LB\n\n## Introduction\nCongratulation to all participants. To begin with, our team would like to thank organizers and Kaggle team for making this competition possible. I also would like to express my gratitude to my teammates: @urvishp80 , @whurobin , @philipkd , and especially to @naoism : he helped me today until 3:30 AM of his time to assemble the final submissions. Below I'll share the key point of our approach, and @naoism may publish a separate post describing specific details of his part.\nThis competition was relatively short for me:  I had just ~2 weeks after finishing another competition. Below I provide the key points of my portion of our solution.\n\n## Model\nMy part of the solution is largely influenced by the method [I proposed in 2020 Cornell Birdcall Identification challene](https://www.kaggle.com/c/birdsong-recognition/discussion/183258) targeted at building a segmentation model based on weak classification labels.\n\n![](https://i.ibb.co/1bZ0Ghg/image.png)\n\nI collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and then apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp (LSE) pooling to produce the prediction global prediction. In this year I quite modified this approach, and the model predictions look much better in comparison with ones from the previous year. First, following [this article](https://arxiv.org/pdf/1411.6228.pdf) I introduced temperature to LSE equal to 5. Another trick is enforcing the minimum of predicted sequence to be 0: bird call is only present in a short time frame. It is achieved by an additional LSE output with -1 temperature followed by applying a loss with zero GT label. These modifications drastically boosted the model performance. Another important thing is using sufficiently short audio segments during training. Long segments help to overcome label weakness, but I realized that a model trained on long segments learns only easy sounds ignoring the rest (it is often just enough to find an easy example in a long sequence to assign the corresponding label). To encourage the model to learn difficult examples I used 5s chunks during training (but I used PL based sampling to deal with weak labels). Important component of the model is a multi-head self-attention applied to entire sequence:  in my experiments this component boosted F1 of sequence based predictions (10 min) by 0.01-0.02 in comparison with CV evaluated based solely on 5s chunks.\n\n**Additional details:**\nBackbone: ResNeXt50\nLoss: Focal loss, corrected to be suitable for soft labels\nAugmentation: [MixUp with max label](https://www.kaggle.com/c/birdsong-recognition/discussion/183199) (proposed by @theoviel ), pink noise, external noise\n**[important]** n_ftt = 2048. It appeared that models trained with n_ftt = 1024 are very susceptible to the domain shift, and n_ftt = 2048 leads to much better performance on soundscapes.\n\n\n## Train-Segment-Shift\nIt is the concept that was used to address the competition problem. **There are 2 major issues to address in this challenge: (1) label weakness and noise, and (2) domain mismatch** between train and test data. My approach consists of 3 steps: train, segment, and shift.\n\n(1) I begin with creating a segmentation model based on the original labels assigned to entire clips. The main difficulty is that clip labels are not equal to chunk labels because the clips may have no bird calls in particular parts or bird calls from other species. So **direct training of a model on clip labels assigned to chunks cannot produce a good model**, but it gives PL for the next step.\n\n(2) Run similar setup as step 1 but do sampling based on PL segmentation, making sure that the primary label bird call is included in each selected audio segment. In addition, I applied segmentation loss to region having high confidence pseudo labels, while masking regions where confidence is not sufficient. It is done to guide the model, and it improved the quality of segmentation labels at the end. \nThe image below shows the model predictions from the same audio for top 5 classes. Blue color corresponds to the primary label, while orange corresponds to the secondary label. The bottom image shows an example of segmentation PL used at step2 and step3. Red bars indicate high confidence PL, while orange ones correspond to low confidence PL (used for masking).\n![](https://i.ibb.co/SXqcYw6/image.png)\nAt this step I got ~0.85 F1 CV (based on short audio) evaluated by taking max values for predictions within entire clip train files (so the provided weak labels can be used for model evaluation). However, if I try to apply this model directly to train soundscapes (which are similar to test data), I get only 0.67 CV (soundscapes), which indicated the domain mismatch between short audio and soundscapes.\n\n(3) That's why the next step is crucial: domain shift accommodation. In this competition the mismatch is not as bad as one year ago (in that case soundscapes were recorded with 16 kHz rate only). Though, if you plot the spectrograms or listen soundscapes, they are still drastically different from short train clips. \nInitially I tried to add noise extracted from train soundscapes (with proper CV split) to each loaded audio segment (I didn't use train soundscapes directly, only noise from them) + pink noise. It didn't work well giving 0.7231/0.6401 at public/private LB. Next, I realized that using [Rain Forest data](https://www.kaggle.com/c/rfcx-species-audio-detection/data) works really well to eliminate FP. This data contains multiple background sounds and sounds similar to bird calls but not bird calls. **So the model learns to ignore sounds similar to bird calls as well as bird calls from unknown species.** This approach boosted the single model performance to  0.7598/0.6486 at public/private LB.\n\nFor example, the image below shows a single model single fold predictions for the first two minutes of 10534_SSW_20170429.ogg soundscape for top 10 classes. The predictions look really confident for detecting nocalls. There is some confusion between some classes though. Also, I noticed that GT labels missed several calls: it is quite clear that the calls are present, but probably experts were not confident enough about the species because of the noise.\n \n![](https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png)\nOne may ask, why do we need step2 and high accuracy segmentation PL. Without noise the model is able to localize the areas with corresponding bird calls. Meanwhile if extensive noise is added, the model needs hints on localization of birdcalls and ignoring the rest. Segmentation PL help to guide training.\n\n\n## Postprocessing\n\nOne important portion of our solution is postprocessing. We averaged the prediction for a given 5s chunk with predictions for chunks shifted by 2.5 seconds forward and backward (one can consider a shifted audio), so `p = 0.5*p0 + 0.25*p_r + 0.25*pl`. This postprocessing is especially effective to deal with cases when a birdcall is located at a chunk boundary. In addition, @naoism was using a postprocessing based on multiplying all chunk predictions for a given class by 1.3 if the maximum predicted value for the given class over the entire sequence is reaching a particular threshold.\n\n## Ensembling\n\nOur best final submission is combination of 2 models having 0.7471/0.6206 and 0.7598/0.6486 public LB and private LB scores which gives 0.7736/0.6605 for weighted ensemble.",
      "votes": null
    },
    {
      "id": "1332458",
      "postDate": "06/02/2021 06:16:08",
      "content": "<p>Congrats on the silver guys. It was great working with you. Learned a lot from <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>.</p>",
      "rawMarkdown": "Congrats on the silver guys. It was great working with you. Learned a lot from @iafoss and @naoism.",
      "votes": null
    },
    {
      "id": "1332468",
      "postDate": "06/02/2021 06:20:43",
      "content": "<p>Thanks so much, I'm glad that we worked together here. It was short but quite intense competition for me. </p>",
      "rawMarkdown": "Thanks so much, I'm glad that we worked together here. It was short but quite intense competition for me.",
      "votes": null
    },
    {
      "id": "1332481",
      "postDate": "06/02/2021 06:30:20",
      "content": "<p>Thanks for sharing your solution. I really like the different graphs. 👍</p>",
      "rawMarkdown": "Thanks for sharing your solution. I really like the different graphs. 👍",
      "votes": null
    },
    {
      "id": "1332491",
      "postDate": "06/02/2021 06:37:18",
      "content": "<p>You are very welcome</p>",
      "rawMarkdown": "You are very welcome",
      "votes": null
    },
    {
      "id": "1332527",
      "postDate": "06/02/2021 06:58:52",
      "content": "<p>Thank you for mentioning my name.<br>\nIt was fun to make final submission late at night, so It's fine.<br>\nI really enjoyed the discussions with the team <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a>, <a href=\"https://www.kaggle.com/whurobin\" target=\"_blank\">@whurobin</a>, <a href=\"https://www.kaggle.com/philipkd\" target=\"_blank\">@philipkd</a>, especially you, and it motivated me even more to work on kaggle. <br>\nThank you very much.</p>",
      "rawMarkdown": "Thank you for mentioning my name.\nIt was fun to make final submission late at night, so It's fine.\nI really enjoyed the discussions with the team @iafoss, @urvishp80, @whurobin, @philipkd, especially you, and it motivated me even more to work on kaggle. \nThank you very much.",
      "votes": null
    },
    {
      "id": "1332784",
      "postDate": "06/02/2021 09:50:48",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and team!! </p>\n<p>ps: At the starting of the competition I've tried many many things based on your previous year pipeline (LSE pooling) but couldn't make it work better than other models at the end. Now I see what I missed. Need to read more carefully your analysis to digest it though. </p>\n<blockquote>\n  <p>I used 5s chunks during training (but I used PL based sampling to deal with weak labels).</p>\n</blockquote>\n<p>Q1) This describes stage 2, right? If I got it correct the PL model is trained on whole clips (?)  <br>\nQ2) single channel with 128 mels I guess <br>\nQ3) Have you tried other attention heads or without (plain fc head) ?</p>",
      "rawMarkdown": "Congrats @iafoss and team!! \n\nps: At the starting of the competition I've tried many many things based on your previous year pipeline (LSE pooling) but couldn't make it work better than other models at the end. Now I see what I missed. Need to read more carefully your analysis to digest it though. \n\n\n> I used 5s chunks during training (but I used PL based sampling to deal with weak labels).\n\nQ1) This describes stage 2, right? If I got it correct the PL model is trained on whole clips (?)  \nQ2) single channel with 128 mels I guess \nQ3) Have you tried other attention heads or without (plain fc head) ?",
      "votes": null
    },
    {
      "id": "1333071",
      "postDate": "06/02/2021 13:23:37",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and team. Good job! </p>",
      "rawMarkdown": "Congrats @iafoss and team. Good job!",
      "votes": null
    },
    {
      "id": "1333096",
      "postDate": "06/02/2021 13:37:41",
      "content": "<p>Thanks.<br>\n(1) Actually, as I mentioned training on longer sequences helps to deal with weak labels but degraded the model ability to handle difficult cases (balance is important). Training on longer clips was helpful in 2020, but here I tried to incorporate the things related to loss (quite helped to improve the model - you can check my previous year plots) and sampling based on PL. I didn't have time to experiment with longer sequences used for training (there were just a few attempts where I tried it without any noticeable improvement). However, longer sequences were very helpful at inference (given the model architecture I use). For example, my model got 0.72-0.73 CV on 5s chunks. While if I do the same validation on 10 min clips passed to the model I got 0.75-0.76, as I remember, and shift based postprocessing boosted the things to ~0.78.<br>\n(2) Yes, it's the same as in the previous year solution. The important thing, though, is using nfft=2048 - it is one of the key points of elimination the domain gap (though private LB performance was similar for nfft=1024 and nfft=2048)<br>\n(3) The time was very short for me. So I just took the model I built in 2020. The thing I'd change or wanted to experiment with? Probably use of a SWIN like transformer base block in the head or attention that incorporates relative positions within fixed window.</p>",
      "rawMarkdown": "Thanks.\n(1) Actually, as I mentioned training on longer sequences helps to deal with weak labels but degraded the model ability to handle difficult cases (balance is important). Training on longer clips was helpful in 2020, but here I tried to incorporate the things related to loss (quite helped to improve the model - you can check my previous year plots) and sampling based on PL. I didn't have time to experiment with longer sequences used for training (there were just a few attempts where I tried it without any noticeable improvement). However, longer sequences were very helpful at inference (given the model architecture I use). For example, my model got 0.72-0.73 CV on 5s chunks. While if I do the same validation on 10 min clips passed to the model I got 0.75-0.76, as I remember, and shift based postprocessing boosted the things to ~0.78.\n(2) Yes, it's the same as in the previous year solution. The important thing, though, is using nfft=2048 - it is one of the key points of elimination the domain gap (though private LB performance was similar for nfft=1024 and nfft=2048)\n(3) The time was very short for me. So I just took the model I built in 2020. The thing I'd change or wanted to experiment with? Probably use of a SWIN like transformer base block in the head or attention that incorporates relative positions within fixed window.",
      "votes": null
    },
    {
      "id": "1333098",
      "postDate": "06/02/2021 13:38:13",
      "content": "<p>Congrats for having a competitive model starting directly with sound signal.  I'll reread this writeup for sure.</p>",
      "rawMarkdown": "Congrats for having a competitive model starting directly with sound signal.  I'll reread this writeup for sure.",
      "votes": null
    },
    {
      "id": "1333129",
      "postDate": "06/02/2021 13:53:34",
      "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. Congratulations to you as well with getting solo gold! I hope you are not very disappointed by the results, but I was really impressed by the score you had during the competition. It motivated me to look into more things. </p>\n<p>Regarding the model, it may be a misunderstanding, but I worked with mel spectrograms as well…  (I think u may get confused with my distribution of predicted labels over time). I approached the task as a segmentation problem and tried to improve my 2020 method as much as I could (I didn't use SED because in 2020 it didn't work well for me). </p>",
      "rawMarkdown": "Thanks so much @cpmpml. Congratulations to you as well with getting solo gold! I hope you are not very disappointed by the results, but I was really impressed by the score you had during the competition. It motivated me to look into more things. \n\nRegarding the model, it may be a misunderstanding, but I worked with mel spectrograms as well...  (I think u may get confused with my distribution of predicted labels over time). I approached the task as a segmentation problem and tried to improve my 2020 method as much as I could (I didn't use SED because in 2020 it didn't work well for me).",
      "votes": null
    },
    {
      "id": "1333131",
      "postDate": "06/02/2021 13:54:13",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a></p>",
      "rawMarkdown": "Thanks @duykhanh99",
      "votes": null
    },
    {
      "id": "1333274",
      "postDate": "06/02/2021 15:42:34",
      "content": "<p>Thanks for your detailed reply. </p>\n<p>PS: Looking more carefully, I see form your plots that we use same mel-spec params resulting in mels 128x584 (or smth close) </p>",
      "rawMarkdown": "Thanks for your detailed reply. \n\nPS: Looking more carefully, I see form your plots that we use same mel-spec params resulting in mels 128x584 (or smth close)",
      "votes": null
    },
    {
      "id": "1333276",
      "postDate": "06/02/2021 15:44:17",
      "content": "<p>LOL, sorry I misread, I am tired.  Still, your model is different from others, and I am sure that including it in any other team ensemble would improve it.  Thanks for the kind words.  I am not too disappointed because I learned a lot about non CNN vision models along the way.</p>",
      "rawMarkdown": "LOL, sorry I misread, I am tired.  Still, your model is different from others, and I am sure that including it in any other team ensemble would improve it.  Thanks for the kind words.  I am not too disappointed because I learned a lot about non CNN vision models along the way.",
      "votes": null
    },
    {
      "id": "1333328",
      "postDate": "06/02/2021 16:24:54",
      "content": "<p>For 5s chunks I have 128x512 size (I set hop length to ensure this dim).</p>",
      "rawMarkdown": "For 5s chunks I have 128x512 size (I set hop length to ensure this dim).",
      "votes": null
    },
    {
      "id": "1334539",
      "postDate": "06/03/2021 15:24:47",
      "content": "<p>Congratz !</p>\n<p>It turns out using Mixup with max label had been used <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">2 years ago</a> by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> - I wasn't aware of that last year, so I'm just sharing the credit  :)</p>",
      "rawMarkdown": "Congratz !\n\nIt turns out using Mixup with max label had been used [2 years ago](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926) by @ddanevskyi - I wasn't aware of that last year, so I'm just sharing the credit  :)",
      "votes": null
    },
    {
      "id": "1334591",
      "postDate": "06/03/2021 15:59:10",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> . I think your previous year solution still was the most influential, at least for our team.</p>",
      "rawMarkdown": "Thanks @theoviel . I think your previous year solution still was the most influential, at least for our team.",
      "votes": null
    },
    {
      "id": "1334627",
      "postDate": "06/03/2021 16:22:33",
      "content": "<p>Glad to hear that, I saw <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> 's post as well :)</p>",
      "rawMarkdown": "Glad to hear that, I saw @naoism 's post as well :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1332458,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "06/02/2021 06:16:08",
      "content": "<p>Congrats on the silver guys. It was great working with you. Learned a lot from <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332468,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/02/2021 06:20:43",
          "content": "<p>Thanks so much, I'm glad that we worked together here. It was short but quite intense competition for me. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332481,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/02/2021 06:30:20",
      "content": "<p>Thanks for sharing your solution. I really like the different graphs. 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332491,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/02/2021 06:37:18",
          "content": "<p>You are very welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332527,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "06/02/2021 06:58:52",
      "content": "<p>Thank you for mentioning my name.<br>\nIt was fun to make final submission late at night, so It's fine.<br>\nI really enjoyed the discussions with the team <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a>, <a href=\"https://www.kaggle.com/whurobin\" target=\"_blank\">@whurobin</a>, <a href=\"https://www.kaggle.com/philipkd\" target=\"_blank\">@philipkd</a>, especially you, and it motivated me even more to work on kaggle. <br>\nThank you very much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332784,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "06/02/2021 09:50:48",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and team!! </p>\n<p>ps: At the starting of the competition I've tried many many things based on your previous year pipeline (LSE pooling) but couldn't make it work better than other models at the end. Now I see what I missed. Need to read more carefully your analysis to digest it though. </p>\n<blockquote>\n  <p>I used 5s chunks during training (but I used PL based sampling to deal with weak labels).</p>\n</blockquote>\n<p>Q1) This describes stage 2, right? If I got it correct the PL model is trained on whole clips (?)  <br>\nQ2) single channel with 128 mels I guess <br>\nQ3) Have you tried other attention heads or without (plain fc head) ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333096,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/02/2021 13:37:41",
          "content": "<p>Thanks.<br>\n(1) Actually, as I mentioned training on longer sequences helps to deal with weak labels but degraded the model ability to handle difficult cases (balance is important). Training on longer clips was helpful in 2020, but here I tried to incorporate the things related to loss (quite helped to improve the model - you can check my previous year plots) and sampling based on PL. I didn't have time to experiment with longer sequences used for training (there were just a few attempts where I tried it without any noticeable improvement). However, longer sequences were very helpful at inference (given the model architecture I use). For example, my model got 0.72-0.73 CV on 5s chunks. While if I do the same validation on 10 min clips passed to the model I got 0.75-0.76, as I remember, and shift based postprocessing boosted the things to ~0.78.<br>\n(2) Yes, it's the same as in the previous year solution. The important thing, though, is using nfft=2048 - it is one of the key points of elimination the domain gap (though private LB performance was similar for nfft=1024 and nfft=2048)<br>\n(3) The time was very short for me. So I just took the model I built in 2020. The thing I'd change or wanted to experiment with? Probably use of a SWIN like transformer base block in the head or attention that incorporates relative positions within fixed window.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333274,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "06/02/2021 15:42:34",
          "content": "<p>Thanks for your detailed reply. </p>\n<p>PS: Looking more carefully, I see form your plots that we use same mel-spec params resulting in mels 128x584 (or smth close) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333328,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/02/2021 16:24:54",
          "content": "<p>For 5s chunks I have 128x512 size (I set hop length to ensure this dim).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333071,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "06/02/2021 13:23:37",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and team. Good job! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1333131,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/02/2021 13:54:13",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333098,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/02/2021 13:38:13",
      "content": "<p>Congrats for having a competitive model starting directly with sound signal.  I'll reread this writeup for sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333129,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/02/2021 13:53:34",
          "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. Congratulations to you as well with getting solo gold! I hope you are not very disappointed by the results, but I was really impressed by the score you had during the competition. It motivated me to look into more things. </p>\n<p>Regarding the model, it may be a misunderstanding, but I worked with mel spectrograms as well…  (I think u may get confused with my distribution of predicted labels over time). I approached the task as a segmentation problem and tried to improve my 2020 method as much as I could (I didn't use SED because in 2020 it didn't work well for me). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333276,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/02/2021 15:44:17",
          "content": "<p>LOL, sorry I misread, I am tired.  Still, your model is different from others, and I am sure that including it in any other team ensemble would improve it.  Thanks for the kind words.  I am not too disappointed because I learned a lot about non CNN vision models along the way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334539,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "06/03/2021 15:24:47",
      "content": "<p>Congratz !</p>\n<p>It turns out using Mixup with max label had been used <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">2 years ago</a> by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> - I wasn't aware of that last year, so I'm just sharing the credit  :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1334591,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/03/2021 15:59:10",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> . I think your previous year solution still was the most influential, at least for our team.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334627,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "06/03/2021 16:22:33",
          "content": "<p>Glad to hear that, I saw <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> 's post as well :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1332447": "## Summary\n- Sequence based model trained on weak labels\n- Multi-head self-attention applied to entire sequences\n- Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.\n- Rainforest noise\n- postprocessing based on 2.5s shifted sequence\n- best **single model performance 0.7598/0.6486** at public/private LB\n\n## Introduction\nCongratulation to all participants. To begin with, our team would like to thank organizers and Kaggle team for making this competition possible. I also would like to express my gratitude to my teammates: @urvishp80 , @whurobin , @philipkd , and especially to @naoism : he helped me today until 3:30 AM of his time to assemble the final submissions. Below I'll share the key point of our approach, and @naoism may publish a separate post describing specific details of his part.\nThis competition was relatively short for me:  I had just ~2 weeks after finishing another competition. Below I provide the key points of my portion of our solution.\n\n## Model\nMy part of the solution is largely influenced by the method [I proposed in 2020 Cornell Birdcall Identification challene](https://www.kaggle.com/c/birdsong-recognition/discussion/183258) targeted at building a segmentation model based on weak classification labels.\n\n![](https://i.ibb.co/1bZ0Ghg/image.png)\n\nI collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and then apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp (LSE) pooling to produce the prediction global prediction. In this year I quite modified this approach, and the model predictions look much better in comparison with ones from the previous year. First, following [this article](https://arxiv.org/pdf/1411.6228.pdf) I introduced temperature to LSE equal to 5. Another trick is enforcing the minimum of predicted sequence to be 0: bird call is only present in a short time frame. It is achieved by an additional LSE output with -1 temperature followed by applying a loss with zero GT label. These modifications drastically boosted the model performance. Another important thing is using sufficiently short audio segments during training. Long segments help to overcome label weakness, but I realized that a model trained on long segments learns only easy sounds ignoring the rest (it is often just enough to find an easy example in a long sequence to assign the corresponding label). To encourage the model to learn difficult examples I used 5s chunks during training (but I used PL based sampling to deal with weak labels). Important component of the model is a multi-head self-attention applied to entire sequence:  in my experiments this component boosted F1 of sequence based predictions (10 min) by 0.01-0.02 in comparison with CV evaluated based solely on 5s chunks.\n\n**Additional details:**\nBackbone: ResNeXt50\nLoss: Focal loss, corrected to be suitable for soft labels\nAugmentation: [MixUp with max label](https://www.kaggle.com/c/birdsong-recognition/discussion/183199) (proposed by @theoviel ), pink noise, external noise\n**[important]** n_ftt = 2048. It appeared that models trained with n_ftt = 1024 are very susceptible to the domain shift, and n_ftt = 2048 leads to much better performance on soundscapes.\n\n\n## Train-Segment-Shift\nIt is the concept that was used to address the competition problem. **There are 2 major issues to address in this challenge: (1) label weakness and noise, and (2) domain mismatch** between train and test data. My approach consists of 3 steps: train, segment, and shift.\n\n(1) I begin with creating a segmentation model based on the original labels assigned to entire clips. The main difficulty is that clip labels are not equal to chunk labels because the clips may have no bird calls in particular parts or bird calls from other species. So **direct training of a model on clip labels assigned to chunks cannot produce a good model**, but it gives PL for the next step.\n\n(2) Run similar setup as step 1 but do sampling based on PL segmentation, making sure that the primary label bird call is included in each selected audio segment. In addition, I applied segmentation loss to region having high confidence pseudo labels, while masking regions where confidence is not sufficient. It is done to guide the model, and it improved the quality of segmentation labels at the end. \nThe image below shows the model predictions from the same audio for top 5 classes. Blue color corresponds to the primary label, while orange corresponds to the secondary label. The bottom image shows an example of segmentation PL used at step2 and step3. Red bars indicate high confidence PL, while orange ones correspond to low confidence PL (used for masking).\n![](https://i.ibb.co/SXqcYw6/image.png)\nAt this step I got ~0.85 F1 CV (based on short audio) evaluated by taking max values for predictions within entire clip train files (so the provided weak labels can be used for model evaluation). However, if I try to apply this model directly to train soundscapes (which are similar to test data), I get only 0.67 CV (soundscapes), which indicated the domain mismatch between short audio and soundscapes.\n\n(3) That's why the next step is crucial: domain shift accommodation. In this competition the mismatch is not as bad as one year ago (in that case soundscapes were recorded with 16 kHz rate only). Though, if you plot the spectrograms or listen soundscapes, they are still drastically different from short train clips. \nInitially I tried to add noise extracted from train soundscapes (with proper CV split) to each loaded audio segment (I didn't use train soundscapes directly, only noise from them) + pink noise. It didn't work well giving 0.7231/0.6401 at public/private LB. Next, I realized that using [Rain Forest data](https://www.kaggle.com/c/rfcx-species-audio-detection/data) works really well to eliminate FP. This data contains multiple background sounds and sounds similar to bird calls but not bird calls. **So the model learns to ignore sounds similar to bird calls as well as bird calls from unknown species.** This approach boosted the single model performance to  0.7598/0.6486 at public/private LB.\n\nFor example, the image below shows a single model single fold predictions for the first two minutes of 10534_SSW_20170429.ogg soundscape for top 10 classes. The predictions look really confident for detecting nocalls. There is some confusion between some classes though. Also, I noticed that GT labels missed several calls: it is quite clear that the calls are present, but probably experts were not confident enough about the species because of the noise.\n \n![](https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png)\nOne may ask, why do we need step2 and high accuracy segmentation PL. Without noise the model is able to localize the areas with corresponding bird calls. Meanwhile if extensive noise is added, the model needs hints on localization of birdcalls and ignoring the rest. Segmentation PL help to guide training.\n\n\n## Postprocessing\n\nOne important portion of our solution is postprocessing. We averaged the prediction for a given 5s chunk with predictions for chunks shifted by 2.5 seconds forward and backward (one can consider a shifted audio), so `p = 0.5*p0 + 0.25*p_r + 0.25*pl`. This postprocessing is especially effective to deal with cases when a birdcall is located at a chunk boundary. In addition, @naoism was using a postprocessing based on multiplying all chunk predictions for a given class by 1.3 if the maximum predicted value for the given class over the entire sequence is reaching a particular threshold.\n\n## Ensembling\n\nOur best final submission is combination of 2 models having 0.7471/0.6206 and 0.7598/0.6486 public LB and private LB scores which gives 0.7736/0.6605 for weighted ensemble.",
    "1332458": "Congrats on the silver guys. It was great working with you. Learned a lot from @iafoss and @naoism.",
    "1332468": "Thanks so much, I'm glad that we worked together here. It was short but quite intense competition for me.",
    "1332481": "Thanks for sharing your solution. I really like the different graphs. 👍",
    "1332491": "You are very welcome",
    "1332527": "Thank you for mentioning my name.\nIt was fun to make final submission late at night, so It's fine.\nI really enjoyed the discussions with the team @iafoss, @urvishp80, @whurobin, @philipkd, especially you, and it motivated me even more to work on kaggle. \nThank you very much.",
    "1332784": "Congrats @iafoss and team!! \n\nps: At the starting of the competition I've tried many many things based on your previous year pipeline (LSE pooling) but couldn't make it work better than other models at the end. Now I see what I missed. Need to read more carefully your analysis to digest it though. \n\n\n> I used 5s chunks during training (but I used PL based sampling to deal with weak labels).\n\nQ1) This describes stage 2, right? If I got it correct the PL model is trained on whole clips (?)  \nQ2) single channel with 128 mels I guess \nQ3) Have you tried other attention heads or without (plain fc head) ?",
    "1333071": "Congrats @iafoss and team. Good job!",
    "1333096": "Thanks.\n(1) Actually, as I mentioned training on longer sequences helps to deal with weak labels but degraded the model ability to handle difficult cases (balance is important). Training on longer clips was helpful in 2020, but here I tried to incorporate the things related to loss (quite helped to improve the model - you can check my previous year plots) and sampling based on PL. I didn't have time to experiment with longer sequences used for training (there were just a few attempts where I tried it without any noticeable improvement). However, longer sequences were very helpful at inference (given the model architecture I use). For example, my model got 0.72-0.73 CV on 5s chunks. While if I do the same validation on 10 min clips passed to the model I got 0.75-0.76, as I remember, and shift based postprocessing boosted the things to ~0.78.\n(2) Yes, it's the same as in the previous year solution. The important thing, though, is using nfft=2048 - it is one of the key points of elimination the domain gap (though private LB performance was similar for nfft=1024 and nfft=2048)\n(3) The time was very short for me. So I just took the model I built in 2020. The thing I'd change or wanted to experiment with? Probably use of a SWIN like transformer base block in the head or attention that incorporates relative positions within fixed window.",
    "1333098": "Congrats for having a competitive model starting directly with sound signal.  I'll reread this writeup for sure.",
    "1333129": "Thanks so much @cpmpml. Congratulations to you as well with getting solo gold! I hope you are not very disappointed by the results, but I was really impressed by the score you had during the competition. It motivated me to look into more things. \n\nRegarding the model, it may be a misunderstanding, but I worked with mel spectrograms as well...  (I think u may get confused with my distribution of predicted labels over time). I approached the task as a segmentation problem and tried to improve my 2020 method as much as I could (I didn't use SED because in 2020 it didn't work well for me).",
    "1333131": "Thanks @duykhanh99",
    "1333274": "Thanks for your detailed reply. \n\nPS: Looking more carefully, I see form your plots that we use same mel-spec params resulting in mels 128x584 (or smth close)",
    "1333276": "LOL, sorry I misread, I am tired.  Still, your model is different from others, and I am sure that including it in any other team ensemble would improve it.  Thanks for the kind words.  I am not too disappointed because I learned a lot about non CNN vision models along the way.",
    "1333328": "For 5s chunks I have 128x512 size (I set hop length to ensure this dim).",
    "1334539": "Congratz !\n\nIt turns out using Mixup with max label had been used [2 years ago](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926) by @ddanevskyi - I wasn't aware of that last year, so I'm just sharing the credit  :)",
    "1334591": "Thanks @theoviel . I think your previous year solution still was the most influential, at least for our team.",
    "1334627": "Glad to hear that, I saw @naoism 's post as well :)"
  },
  "source": "meta"
}