{
  "id": 220342,
  "title": "4th place solution (5th public LB)",
  "url": "/competitions/rfcx-species-audio-detection/writeups/hot-pursuit-4th-place-solution-5th-public-lb",
  "author_name": "",
  "post_date": "2021-02-18T07:13:59.773Z",
  "votes": 61,
  "comment_count": 16,
  "views": 0,
  "content": "<p>First of all, I'd like to thank my teammates, Rainforest Connection and Kaggle for this interesting and tricky challenge!</p>\n<p>The major issue in this competition was obviously the labelling quality. True- and False-Positives audios contain lots of unlabelled regions that adds too much noise for the models. As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.</p>\n<p>Moreover, labeled regions in the TP audios were balanced, i.e. each class had an equal number of labels. However, we've noticed that test predictions contain mostly the 3rd class as a top-1 probability. And with the higher percentage of the 3rd class, the LB score tends to be better. The similar situation was for some other classes (e.g. top-2 was mostly the 18th class). It gave us an idea that probably test files have completely different class distributions compared to the TP data. That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).</p>\n<p>Our final solution consists of 3 stages.</p>\n<h2>1st Stage</h2>\n<ul>\n<li>Data: only TP labels on 26 classes (for each song_type).</li>\n<li>Models: SED-classifiers (EfficientNet-B1 and EfficientNet-B3)</li>\n<li>Cropping strategy: Random crops around TP regions</li>\n<li>Loss: BCE</li>\n<li>Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix: cutting the TP regions and pasting them into the random time regions in the other TP and FP audios.</li>\n<li>Public LB score: 0.909 -&gt; 940 (after balancing)</li>\n<li>Private LB score: 0.915 -&gt; 0.938</li>\n</ul>\n<h2>2nd Stage</h2>\n<p>Taking the models from the 1st Stage we've made a set of pseudolabels for TP (OOF), FP and test data. The pseudolabels have been generated using the SED framewise output. At this point, audio files have much more labeled regions compared to the initial TP data. And on this stage models are being trained on the pseudolabels only. We've applied two approaches:</p>\n<h5>SED-classification</h5>\n<ul>\n<li>Data: TP pseudolabels  +  random 2000 samples from FP pseudolabels for each fold. Use soft labels (0.9) for the pseudolabels</li>\n<li>Models: SED-classifiers (EfficientNet-B0, EfficientNet-B1, MobileNetV2, DenseNet121)</li>\n<li>Cropping strategy: Random 5 seconds crops around pseudolabeled regions</li>\n<li>Loss: modified LSEP loss</li>\n<li>Augmentations: raw audio augmentations, such as: GaussianNoiseSNR, PinkNoiseSNR, PitchShift, TimeShift, VolumeControl </li>\n<li>TTA: 6 different crop sizes are used during the inference: 2, 5, 10, 20, 30 and 60 seconds</li>\n<li>Best single model (5 fold) public LB score: 0.957 (after balancing)</li>\n<li>Private LB score: 0.963</li>\n</ul>\n<h5>Usual classification</h5>\n<ul>\n<li>Data: TP + FP pseudolabels. Pre-train models on the test pseudolabels</li>\n<li>Models: Usual classifiers (EfficientNet-B1, ResNet34, SE-ResNeXt50)</li>\n<li>Cropping strategy: Random crops around pseudolabeled regions</li>\n<li>Loss: BCE</li>\n<li>Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix</li>\n<li>Best single model (5 fold) public LB score: 0.952 (after balancing)</li>\n<li>Private LB score: 0.959</li>\n</ul>\n<h2>3rd Stage</h2>\n<p>Taking the overall ensemble from the 2nd Stage allows to get the Public LB score of 0.965 (Private LB: 0.969). To achieve our best 0.969 Public LB (Private LB: 0.971) we're applying single class semantic segmentation models for 3rd, 11th and 18th classes (other classes didn't give any score improvements on the Public LB).</p>\n<p>The segmentation polish is done in the following manner:<br>\n<code>class_score = class_score * (1 + 0.1 * num_instances) if num_instances &gt; 0 else  class_score * 0.9</code>, where <code>num_instances</code> is the number of instances predicted by the semantic segmentation model for each recording.</p>\n<h2>What didn`t work</h2>\n<ol>\n<li>PANN pretrained weights (or other audio pretrained models) - imagenet performs best</li>\n<li>Using \"fat\" encoders</li>\n<li>Focal loss with soft penalty (But as we see It works for the other participants)</li>\n<li>Multiclass segmentation</li>\n<li>Raw audio classification with 1d convolutions </li>\n</ol>",
  "messages": [
    {
      "id": "1207813",
      "postDate": "02/18/2021 03:48:05",
      "content": "<p>First of all, I'd like to thank my teammates, Rainforest Connection and Kaggle for this interesting and tricky challenge!</p>\n<p>The major issue in this competition was obviously the labelling quality. True- and False-Positives audios contain lots of unlabelled regions that adds too much noise for the models. As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.</p>\n<p>Moreover, labeled regions in the TP audios were balanced, i.e. each class had an equal number of labels. However, we've noticed that test predictions contain mostly the 3rd class as a top-1 probability. And with the higher percentage of the 3rd class, the LB score tends to be better. The similar situation was for some other classes (e.g. top-2 was mostly the 18th class). It gave us an idea that probably test files have completely different class distributions compared to the TP data. That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).</p>\n<p>Our final solution consists of 3 stages.</p>\n<h2>1st Stage</h2>\n<ul>\n<li>Data: only TP labels on 26 classes (for each song_type).</li>\n<li>Models: SED-classifiers (EfficientNet-B1 and EfficientNet-B3)</li>\n<li>Cropping strategy: Random crops around TP regions</li>\n<li>Loss: BCE</li>\n<li>Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix: cutting the TP regions and pasting them into the random time regions in the other TP and FP audios.</li>\n<li>Public LB score: 0.909 -&gt; 940 (after balancing)</li>\n<li>Private LB score: 0.915 -&gt; 0.938</li>\n</ul>\n<h2>2nd Stage</h2>\n<p>Taking the models from the 1st Stage we've made a set of pseudolabels for TP (OOF), FP and test data. The pseudolabels have been generated using the SED framewise output. At this point, audio files have much more labeled regions compared to the initial TP data. And on this stage models are being trained on the pseudolabels only. We've applied two approaches:</p>\n<h5>SED-classification</h5>\n<ul>\n<li>Data: TP pseudolabels  +  random 2000 samples from FP pseudolabels for each fold. Use soft labels (0.9) for the pseudolabels</li>\n<li>Models: SED-classifiers (EfficientNet-B0, EfficientNet-B1, MobileNetV2, DenseNet121)</li>\n<li>Cropping strategy: Random 5 seconds crops around pseudolabeled regions</li>\n<li>Loss: modified LSEP loss</li>\n<li>Augmentations: raw audio augmentations, such as: GaussianNoiseSNR, PinkNoiseSNR, PitchShift, TimeShift, VolumeControl </li>\n<li>TTA: 6 different crop sizes are used during the inference: 2, 5, 10, 20, 30 and 60 seconds</li>\n<li>Best single model (5 fold) public LB score: 0.957 (after balancing)</li>\n<li>Private LB score: 0.963</li>\n</ul>\n<h5>Usual classification</h5>\n<ul>\n<li>Data: TP + FP pseudolabels. Pre-train models on the test pseudolabels</li>\n<li>Models: Usual classifiers (EfficientNet-B1, ResNet34, SE-ResNeXt50)</li>\n<li>Cropping strategy: Random crops around pseudolabeled regions</li>\n<li>Loss: BCE</li>\n<li>Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix</li>\n<li>Best single model (5 fold) public LB score: 0.952 (after balancing)</li>\n<li>Private LB score: 0.959</li>\n</ul>\n<h2>3rd Stage</h2>\n<p>Taking the overall ensemble from the 2nd Stage allows to get the Public LB score of 0.965 (Private LB: 0.969). To achieve our best 0.969 Public LB (Private LB: 0.971) we're applying single class semantic segmentation models for 3rd, 11th and 18th classes (other classes didn't give any score improvements on the Public LB).</p>\n<p>The segmentation polish is done in the following manner:<br>\n<code>class_score = class_score * (1 + 0.1 * num_instances) if num_instances &gt; 0 else  class_score * 0.9</code>, where <code>num_instances</code> is the number of instances predicted by the semantic segmentation model for each recording.</p>\n<h2>What didn`t work</h2>\n<ol>\n<li>PANN pretrained weights (or other audio pretrained models) - imagenet performs best</li>\n<li>Using \"fat\" encoders</li>\n<li>Focal loss with soft penalty (But as we see It works for the other participants)</li>\n<li>Multiclass segmentation</li>\n<li>Raw audio classification with 1d convolutions </li>\n</ol>",
      "rawMarkdown": "First of all, I'd like to thank my teammates, Rainforest Connection and Kaggle for this interesting and tricky challenge!\n\nThe major issue in this competition was obviously the labelling quality. True- and False-Positives audios contain lots of unlabelled regions that adds too much noise for the models. As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.\n\nMoreover, labeled regions in the TP audios were balanced, i.e. each class had an equal number of labels. However, we've noticed that test predictions contain mostly the 3rd class as a top-1 probability. And with the higher percentage of the 3rd class, the LB score tends to be better. The similar situation was for some other classes (e.g. top-2 was mostly the 18th class). It gave us an idea that probably test files have completely different class distributions compared to the TP data. That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).\n\nOur final solution consists of 3 stages.\n\n## 1st Stage\n* Data: only TP labels on 26 classes (for each song_type).\n* Models: SED-classifiers (EfficientNet-B1 and EfficientNet-B3)\n* Cropping strategy: Random crops around TP regions\n* Loss: BCE\n* Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix: cutting the TP regions and pasting them into the random time regions in the other TP and FP audios.\n* Public LB score: 0.909 -> 940 (after balancing)\n* Private LB score: 0.915 -> 0.938\n\n## 2nd Stage\nTaking the models from the 1st Stage we've made a set of pseudolabels for TP (OOF), FP and test data. The pseudolabels have been generated using the SED framewise output. At this point, audio files have much more labeled regions compared to the initial TP data. And on this stage models are being trained on the pseudolabels only. We've applied two approaches:\n\n##### SED-classification\n* Data: TP pseudolabels  +  random 2000 samples from FP pseudolabels for each fold. Use soft labels (0.9) for the pseudolabels\n* Models: SED-classifiers (EfficientNet-B0, EfficientNet-B1, MobileNetV2, DenseNet121)\n* Cropping strategy: Random 5 seconds crops around pseudolabeled regions\n* Loss: modified LSEP loss\n* Augmentations: raw audio augmentations, such as: GaussianNoiseSNR, PinkNoiseSNR, PitchShift, TimeShift, VolumeControl \n* TTA: 6 different crop sizes are used during the inference: 2, 5, 10, 20, 30 and 60 seconds\n* Best single model (5 fold) public LB score: 0.957 (after balancing)\n* Private LB score: 0.963\n\n##### Usual classification\n* Data: TP + FP pseudolabels. Pre-train models on the test pseudolabels\n* Models: Usual classifiers (EfficientNet-B1, ResNet34, SE-ResNeXt50)\n* Cropping strategy: Random crops around pseudolabeled regions\n* Loss: BCE\n* Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix\n* Best single model (5 fold) public LB score: 0.952 (after balancing)\n* Private LB score: 0.959\n\n## 3rd Stage\nTaking the overall ensemble from the 2nd Stage allows to get the Public LB score of 0.965 (Private LB: 0.969). To achieve our best 0.969 Public LB (Private LB: 0.971) we're applying single class semantic segmentation models for 3rd, 11th and 18th classes (other classes didn't give any score improvements on the Public LB).\n\nThe segmentation polish is done in the following manner:\n`class_score = class_score * (1 + 0.1 * num_instances) if num_instances > 0 else  class_score * 0.9`, where `num_instances` is the number of instances predicted by the semantic segmentation model for each recording.\n\n## What didn`t work \n1. PANN pretrained weights (or other audio pretrained models) - imagenet performs best\n2. Using \"fat\" encoders\n3. Focal loss with soft penalty (But as we see It works for the other participants)\n4. Multiclass segmentation\n5. Raw audio classification with 1d convolutions",
      "votes": null
    },
    {
      "id": "1207899",
      "postDate": "02/18/2021 04:53:51",
      "content": "<p>any github link？tks</p>",
      "rawMarkdown": "any github link？tks",
      "votes": null
    },
    {
      "id": "1207930",
      "postDate": "02/18/2021 05:01:32",
      "content": "<p>Congrats and thanks for the writeup <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a> and team</p>",
      "rawMarkdown": "Congrats and thanks for the writeup @kupchanski and team",
      "votes": null
    },
    {
      "id": "1208014",
      "postDate": "02/18/2021 05:48:12",
      "content": "<p>Thanks, interesting to see how class balancing helped with boosting the score. </p>",
      "rawMarkdown": "Thanks, interesting to see how class balancing helped with boosting the score.",
      "votes": null
    },
    {
      "id": "1208016",
      "postDate": "02/18/2021 05:49:53",
      "content": "<p>Thanks for the write-up. Seems that pseudolabeling was quite important to many people to recover from the missing labels. What made you decide to do semantic segmentation for just the few classes?</p>\n<p>We also used a rebalancing factor on a few of the classes and it seemed to be quiet effective. </p>",
      "rawMarkdown": "Thanks for the write-up. Seems that pseudolabeling was quite important to many people to recover from the missing labels. What made you decide to do semantic segmentation for just the few classes?\n\nWe also used a rebalancing factor on a few of the classes and it seemed to be quiet effective.",
      "votes": null
    },
    {
      "id": "1208063",
      "postDate": "02/18/2021 06:25:44",
      "content": "<p>\" As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.\"</p>\n<p>this is a rare competition that you can only rely on smart probing. This is a competition with no ground truth data at all (i.e. clip label which you would use in evaluation metric)</p>\n<hr>\n<p>\"stage.1 Public LB score: 0.909 -&gt; 940 (after balancing)\".</p>\n<p>just for your information, a tp annotation can have actually multiple labels in the time interval tmin to tmax.<br>\ni verify this later by manual inspection.</p>\n<p>hence training tp alone will not give high score. in my experiments, training with tp only is 0.905 (5 fold), training with tp and fp is 0.937</p>",
      "rawMarkdown": "\" As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.\"\n\nthis is a rare competition that you can only rely on smart probing. This is a competition with no ground truth data at all (i.e. clip label which you would use in evaluation metric)\n\n---\n\"stage.1 Public LB score: 0.909 -> 940 (after balancing)\".\n\njust for your information, a tp annotation can have actually multiple labels in the time interval tmin to tmax.\ni verify this later by manual inspection.\n\nhence training tp alone will not give high score. in my experiments, training with tp only is 0.905 (5 fold), training with tp and fp is 0.937",
      "votes": null
    },
    {
      "id": "1208079",
      "postDate": "02/18/2021 06:37:42",
      "content": "<p>Congratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956… This is crazy!</p>",
      "rawMarkdown": "Congratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956... This is crazy!",
      "votes": null
    },
    {
      "id": "1208114",
      "postDate": "02/18/2021 07:00:55",
      "content": "<blockquote>\n  <p>tp annotation can have actually multiple labels in the time interval tmin to tmax.</p>\n</blockquote>\n<p>yeah, so we found all labels that appears on every crop and add them to gt labels for this crop.</p>",
      "rawMarkdown": "> tp annotation can have actually multiple labels in the time interval tmin to tmax.\n\nyeah, so we found all labels that appears on every crop and add them to gt labels for this crop.",
      "votes": null
    },
    {
      "id": "1208131",
      "postDate": "02/18/2021 07:12:03",
      "content": "<p>We tried to teach binary semantic segmentation for every class, but most of them trained poorly with low IOU score on validation ( e.g. 1, 12 classes), others  just decrease LB score  (e.g. 7, 15 classes). So we decide to use only classses which were most frequent to decrease fp quantity.</p>",
      "rawMarkdown": "We tried to teach binary semantic segmentation for every class, but most of them trained poorly with low IOU score on validation ( e.g. 1, 12 classes), others  just decrease LB score  (e.g. 7, 15 classes). So we decide to use only classses which were most frequent to decrease fp quantity.",
      "votes": null
    },
    {
      "id": "1208389",
      "postDate": "02/18/2021 09:26:58",
      "content": "<p>Very nice solution, congratz ! </p>\n<blockquote>\n  <p>That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).</p>\n</blockquote>\n<p>Interesting, we tried this as well but it didn't work. I guess our models already captured these classes well enough.</p>",
      "rawMarkdown": "Very nice solution, congratz ! \n\n> That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).\n\nInteresting, we tried this as well but it didn't work. I guess our models already captured these classes well enough.",
      "votes": null
    },
    {
      "id": "1208475",
      "postDate": "02/18/2021 09:49:41",
      "content": "<p>Congratulations on your success! <br>\nSmooth and clear progress through each stage.</p>",
      "rawMarkdown": "Congratulations on your success! \nSmooth and clear progress through each stage.",
      "votes": null
    },
    {
      "id": "1208660",
      "postDate": "02/18/2021 12:08:18",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "1209901",
      "postDate": "02/19/2021 05:15:10",
      "content": "<p>congrats and thank you for sharing.</p>",
      "rawMarkdown": "congrats and thank you for sharing.",
      "votes": null
    },
    {
      "id": "1210084",
      "postDate": "02/19/2021 07:44:36",
      "content": "<p>Congrats and thanks for sharing.  Interesting that you too made pseudo labeling work.</p>\n<p>A detail question: you say you use random 5 second crops around pseudo label regions, but some regions have a duration of 8 seconds.  You crop them to 5 seconds around their center?</p>",
      "rawMarkdown": "Congrats and thanks for sharing.  Interesting that you too made pseudo labeling work.\n\nA detail question: you say you use random 5 second crops around pseudo label regions, but some regions have a duration of 8 seconds.  You crop them to 5 seconds around their center?",
      "votes": null
    },
    {
      "id": "1210253",
      "postDate": "02/19/2021 09:42:50",
      "content": "<p>Yeah, always 5 sec region crop around center with some random shifts despite the length of event. </p>",
      "rawMarkdown": "Yeah, always 5 sec region crop around center with some random shifts despite the length of event.",
      "votes": null
    },
    {
      "id": "1216198",
      "postDate": "02/24/2021 06:30:20",
      "content": "<p>amazing analysis on tp and fp!!!</p>",
      "rawMarkdown": "amazing analysis on tp and fp!!!",
      "votes": null
    },
    {
      "id": "3095854",
      "postDate": "01/13/2025 20:32:34",
      "content": "<p>Thank you for the solution, well done!</p>",
      "rawMarkdown": "Thank you for the solution, well done!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207899,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "02/18/2021 04:53:51",
      "content": "<p>any github link？tks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207930,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 05:01:32",
      "content": "<p>Congrats and thanks for the writeup <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a> and team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208014,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "02/18/2021 05:48:12",
      "content": "<p>Thanks, interesting to see how class balancing helped with boosting the score. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208016,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 05:49:53",
      "content": "<p>Thanks for the write-up. Seems that pseudolabeling was quite important to many people to recover from the missing labels. What made you decide to do semantic segmentation for just the few classes?</p>\n<p>We also used a rebalancing factor on a few of the classes and it seemed to be quiet effective. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1208131,
          "author_name": "kupchanski",
          "author_url": "",
          "post_date": "02/18/2021 07:12:03",
          "content": "<p>We tried to teach binary semantic segmentation for every class, but most of them trained poorly with low IOU score on validation ( e.g. 1, 12 classes), others  just decrease LB score  (e.g. 7, 15 classes). So we decide to use only classses which were most frequent to decrease fp quantity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208063,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 06:25:44",
      "content": "<p>\" As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.\"</p>\n<p>this is a rare competition that you can only rely on smart probing. This is a competition with no ground truth data at all (i.e. clip label which you would use in evaluation metric)</p>\n<hr>\n<p>\"stage.1 Public LB score: 0.909 -&gt; 940 (after balancing)\".</p>\n<p>just for your information, a tp annotation can have actually multiple labels in the time interval tmin to tmax.<br>\ni verify this later by manual inspection.</p>\n<p>hence training tp alone will not give high score. in my experiments, training with tp only is 0.905 (5 fold), training with tp and fp is 0.937</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208114,
          "author_name": "kupchanski",
          "author_url": "",
          "post_date": "02/18/2021 07:00:55",
          "content": "<blockquote>\n  <p>tp annotation can have actually multiple labels in the time interval tmin to tmax.</p>\n</blockquote>\n<p>yeah, so we found all labels that appears on every crop and add them to gt labels for this crop.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1216198,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "02/24/2021 06:30:20",
          "content": "<p>amazing analysis on tp and fp!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208079,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "02/18/2021 06:37:42",
      "content": "<p>Congratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956… This is crazy!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208389,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 09:26:58",
      "content": "<p>Very nice solution, congratz ! </p>\n<blockquote>\n  <p>That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).</p>\n</blockquote>\n<p>Interesting, we tried this as well but it didn't work. I guess our models already captured these classes well enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208475,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 09:49:41",
      "content": "<p>Congratulations on your success! <br>\nSmooth and clear progress through each stage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208660,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "02/18/2021 12:08:18",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209901,
      "author_name": "harip98",
      "author_url": "",
      "post_date": "02/19/2021 05:15:10",
      "content": "<p>congrats and thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210084,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/19/2021 07:44:36",
      "content": "<p>Congrats and thanks for sharing.  Interesting that you too made pseudo labeling work.</p>\n<p>A detail question: you say you use random 5 second crops around pseudo label regions, but some regions have a duration of 8 seconds.  You crop them to 5 seconds around their center?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210253,
          "author_name": "kupchanski",
          "author_url": "",
          "post_date": "02/19/2021 09:42:50",
          "content": "<p>Yeah, always 5 sec region crop around center with some random shifts despite the length of event. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3095854,
      "author_name": "akhalilo",
      "author_url": "",
      "post_date": "01/13/2025 20:32:34",
      "content": "<p>Thank you for the solution, well done!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207813": "First of all, I'd like to thank my teammates, Rainforest Connection and Kaggle for this interesting and tricky challenge!\n\nThe major issue in this competition was obviously the labelling quality. True- and False-Positives audios contain lots of unlabelled regions that adds too much noise for the models. As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.\n\nMoreover, labeled regions in the TP audios were balanced, i.e. each class had an equal number of labels. However, we've noticed that test predictions contain mostly the 3rd class as a top-1 probability. And with the higher percentage of the 3rd class, the LB score tends to be better. The similar situation was for some other classes (e.g. top-2 was mostly the 18th class). It gave us an idea that probably test files have completely different class distributions compared to the TP data. That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).\n\nOur final solution consists of 3 stages.\n\n## 1st Stage\n* Data: only TP labels on 26 classes (for each song_type).\n* Models: SED-classifiers (EfficientNet-B1 and EfficientNet-B3)\n* Cropping strategy: Random crops around TP regions\n* Loss: BCE\n* Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix: cutting the TP regions and pasting them into the random time regions in the other TP and FP audios.\n* Public LB score: 0.909 -> 940 (after balancing)\n* Private LB score: 0.915 -> 0.938\n\n## 2nd Stage\nTaking the models from the 1st Stage we've made a set of pseudolabels for TP (OOF), FP and test data. The pseudolabels have been generated using the SED framewise output. At this point, audio files have much more labeled regions compared to the initial TP data. And on this stage models are being trained on the pseudolabels only. We've applied two approaches:\n\n##### SED-classification\n* Data: TP pseudolabels  +  random 2000 samples from FP pseudolabels for each fold. Use soft labels (0.9) for the pseudolabels\n* Models: SED-classifiers (EfficientNet-B0, EfficientNet-B1, MobileNetV2, DenseNet121)\n* Cropping strategy: Random 5 seconds crops around pseudolabeled regions\n* Loss: modified LSEP loss\n* Augmentations: raw audio augmentations, such as: GaussianNoiseSNR, PinkNoiseSNR, PitchShift, TimeShift, VolumeControl \n* TTA: 6 different crop sizes are used during the inference: 2, 5, 10, 20, 30 and 60 seconds\n* Best single model (5 fold) public LB score: 0.957 (after balancing)\n* Private LB score: 0.963\n\n##### Usual classification\n* Data: TP + FP pseudolabels. Pre-train models on the test pseudolabels\n* Models: Usual classifiers (EfficientNet-B1, ResNet34, SE-ResNeXt50)\n* Cropping strategy: Random crops around pseudolabeled regions\n* Loss: BCE\n* Augmentations: spectrogram augmentations (SpecAugment, Noise) and CutMix\n* Best single model (5 fold) public LB score: 0.952 (after balancing)\n* Private LB score: 0.959\n\n## 3rd Stage\nTaking the overall ensemble from the 2nd Stage allows to get the Public LB score of 0.965 (Private LB: 0.969). To achieve our best 0.969 Public LB (Private LB: 0.971) we're applying single class semantic segmentation models for 3rd, 11th and 18th classes (other classes didn't give any score improvements on the Public LB).\n\nThe segmentation polish is done in the following manner:\n`class_score = class_score * (1 + 0.1 * num_instances) if num_instances > 0 else  class_score * 0.9`, where `num_instances` is the number of instances predicted by the semantic segmentation model for each recording.\n\n## What didn`t work \n1. PANN pretrained weights (or other audio pretrained models) - imagenet performs best\n2. Using \"fat\" encoders\n3. Focal loss with soft penalty (But as we see It works for the other participants)\n4. Multiclass segmentation\n5. Raw audio classification with 1d convolutions",
    "1207899": "any github link？tks",
    "1207930": "Congrats and thanks for the writeup @kupchanski and team",
    "1208014": "Thanks, interesting to see how class balancing helped with boosting the score.",
    "1208016": "Thanks for the write-up. Seems that pseudolabeling was quite important to many people to recover from the missing labels. What made you decide to do semantic segmentation for just the few classes?\n\nWe also used a rebalancing factor on a few of the classes and it seemed to be quiet effective.",
    "1208063": "\" As a consequence, till the very end of the competition we haven't managed to establish a reliable local validation strategy and were mostly relying on the Public LB scores.\"\n\nthis is a rare competition that you can only rely on smart probing. This is a competition with no ground truth data at all (i.e. clip label which you would use in evaluation metric)\n\n---\n\"stage.1 Public LB score: 0.909 -> 940 (after balancing)\".\n\njust for your information, a tp annotation can have actually multiple labels in the time interval tmin to tmax.\ni verify this later by manual inspection.\n\nhence training tp alone will not give high score. in my experiments, training with tp only is 0.905 (5 fold), training with tp and fp is 0.937",
    "1208079": "Congratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956... This is crazy!",
    "1208114": "> tp annotation can have actually multiple labels in the time interval tmin to tmax.\n\nyeah, so we found all labels that appears on every crop and add them to gt labels for this crop.",
    "1208131": "We tried to teach binary semantic segmentation for every class, but most of them trained poorly with low IOU score on validation ( e.g. 1, 12 classes), others  just decrease LB score  (e.g. 7, 15 classes). So we decide to use only classses which were most frequent to decrease fp quantity.",
    "1208389": "Very nice solution, congratz ! \n\n> That's why we've applied additional multipliers for the 3rd and 18th class to artificially increase probabilities for them (naming it class balancing).\n\nInteresting, we tried this as well but it didn't work. I guess our models already captured these classes well enough.",
    "1208475": "Congratulations on your success! \nSmooth and clear progress through each stage.",
    "1208660": "Congratulations",
    "1209901": "congrats and thank you for sharing.",
    "1210084": "Congrats and thanks for sharing.  Interesting that you too made pseudo labeling work.\n\nA detail question: you say you use random 5 second crops around pseudo label regions, but some regions have a duration of 8 seconds.  You crop them to 5 seconds around their center?",
    "1210253": "Yeah, always 5 sec region crop around center with some random shifts despite the length of event.",
    "1216198": "amazing analysis on tp and fp!!!",
    "3095854": "Thank you for the solution, well done!"
  },
  "source": "meta"
}