{
  "id": 220443,
  "title": "7th place solution - Beluga & Peter",
  "url": "/competitions/rfcx-species-audio-detection/writeups/beluga-peter-7th-place-solution-beluga-peter",
  "author_name": "",
  "post_date": "2021-02-19T08:00:27.723Z",
  "votes": 43,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks, Kaggle and RFCx, for this audio competition, and special thanks to my teammate <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> Without him, I probably would have given up a long time ago, somewhere at 0.8xx.</p>\n<h2>Data preparation</h2>\n<p>Resampling everything to 32kHz and split the audio files into 3 seconds duration chunks. We used a sliding window with a 1-second step.</p>\n<h2>Collecting more label</h2>\n<p>The key to our result was that Beluga collected tons of training samples manually. He created an awesome annotation application; you can find the details <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220305\" target=\"_blank\">here</a>. <em>Source code included</em>.<br>\nAfter the first batch of manually labeled examples, we quickly achieved 0.93x with an ensemble of a varying number of PANN (cnn14) models. </p>\n<h2>Input</h2>\n<p>We used mel-spectrograms as inputs with various <code>n_bin</code> (128, 192, 256, 288). Beluga trained PANN-cnn14 models with one input channel. For the other backbones (effnets, resnets, etc) I used three input channels with a simple trick:</p>\n<ul>\n<li>I used different <code>n_mel</code> and <code>n_fft</code> settings for every channel. E.g. n_mel=(128, 192, 256), n_fft=(1024, 1594, 2048). This results in different height images, so resizing to the same value is necessary.<br>\nWe both used <code>torchlibrosa</code> to generate the mel-spectrograms.</li>\n</ul>\n<h2>Augmentation</h2>\n<p>We used three simple augmentations with different probability:</p>\n<h5>Roll</h5>\n<pre><code>np.roll(y, shift=np.random.randint(0, len(y)))\n</code></pre>\n<h5>Audio mixup</h5>\n<pre><code>w = np.random.uniform(0.3, 0.7)\nmixed = (audio_chunk + rnd_audio_chunk * w) / (1 + w)\nlabel = (label + rnd_label).clip(0, 1)\n</code></pre>\n<h5>Spec augment</h5>\n<pre><code>SpecAugmentation(time_drop_width=16, time_stripes_num=2, freq_drop_width=16, freq_stripes_num=2)\n</code></pre>\n<h2>Architectures</h2>\n<ul>\n<li>PANN - cnn14</li>\n<li>EfficientNet B0, B1, B2</li>\n<li>Densenet 121</li>\n<li>Resnet-50</li>\n<li>Resnest-50</li>\n<li>Mobilnet v3 large 100<br>\nWe trained many versions of these models with different augmentation settings and training data. Beluga used a PANN - cnn14 model (I think it is the same as the original) from his <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183407\" target=\"_blank\">Cornell solution</a>.<br>\nI trained a very similar architecture with different backbones and I used attention head from SED:</li>\n</ul>\n<pre><code>x = ...generate mel-specotrogram...\nx = self.backbone.forward_features(x)\nx = torch.mean(x, dim=2)\nx1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\nx2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\nx = x1 + x2\nx = F.dropout(x, p=0.5, training=self.training)\nx = x.transpose(1, 2)\nx = F.relu_(self.fc1(x))\nx = x.transpose(1, 2)\nx = F.dropout(x, p=0.5, training=self.training)\n(clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\nsegmentwise_output = segmentwise_output.transpose(1, 2)\nframewise_output = interpolate(segmentwise_output, self.interpolate_ratio)\noutput_dict = {\n   \"framewise_output\": framewise_output,\n   \"clipwise_output\": clipwise_output,\n}\n</code></pre>\n<h2>Training</h2>\n<p>Nothing special. Our training method was the same for all of the models:</p>\n<ul>\n<li>4 folds</li>\n<li>10 epochs (15 with a higher probability of mixup)</li>\n<li>Adam (1e-3; *0.9 after 5 epochs)</li>\n<li>BCE loss (PANN version)</li>\n<li>We used the weights of the best validation LWRAP epoch for inference.</li>\n</ul>\n<h2>Pseudo labeling</h2>\n<p>After we had an excellent ensembled score on the public LB (0.950), we started to add pseudo labels to our dataset. The result after we re-trained everything was 0.96x.</p>\n<h2>Final Ensemble</h2>\n<p>Our final ensemble had 80+ models (all of them trained with 4-folds) </p>",
  "messages": [
    {
      "id": "1208483",
      "postDate": "02/18/2021 09:51:22",
      "content": "<p>Thanks, Kaggle and RFCx, for this audio competition, and special thanks to my teammate <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> Without him, I probably would have given up a long time ago, somewhere at 0.8xx.</p>\n<h2>Data preparation</h2>\n<p>Resampling everything to 32kHz and split the audio files into 3 seconds duration chunks. We used a sliding window with a 1-second step.</p>\n<h2>Collecting more label</h2>\n<p>The key to our result was that Beluga collected tons of training samples manually. He created an awesome annotation application; you can find the details <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220305\" target=\"_blank\">here</a>. <em>Source code included</em>.<br>\nAfter the first batch of manually labeled examples, we quickly achieved 0.93x with an ensemble of a varying number of PANN (cnn14) models. </p>\n<h2>Input</h2>\n<p>We used mel-spectrograms as inputs with various <code>n_bin</code> (128, 192, 256, 288). Beluga trained PANN-cnn14 models with one input channel. For the other backbones (effnets, resnets, etc) I used three input channels with a simple trick:</p>\n<ul>\n<li>I used different <code>n_mel</code> and <code>n_fft</code> settings for every channel. E.g. n_mel=(128, 192, 256), n_fft=(1024, 1594, 2048). This results in different height images, so resizing to the same value is necessary.<br>\nWe both used <code>torchlibrosa</code> to generate the mel-spectrograms.</li>\n</ul>\n<h2>Augmentation</h2>\n<p>We used three simple augmentations with different probability:</p>\n<h5>Roll</h5>\n<pre><code>np.roll(y, shift=np.random.randint(0, len(y)))\n</code></pre>\n<h5>Audio mixup</h5>\n<pre><code>w = np.random.uniform(0.3, 0.7)\nmixed = (audio_chunk + rnd_audio_chunk * w) / (1 + w)\nlabel = (label + rnd_label).clip(0, 1)\n</code></pre>\n<h5>Spec augment</h5>\n<pre><code>SpecAugmentation(time_drop_width=16, time_stripes_num=2, freq_drop_width=16, freq_stripes_num=2)\n</code></pre>\n<h2>Architectures</h2>\n<ul>\n<li>PANN - cnn14</li>\n<li>EfficientNet B0, B1, B2</li>\n<li>Densenet 121</li>\n<li>Resnet-50</li>\n<li>Resnest-50</li>\n<li>Mobilnet v3 large 100<br>\nWe trained many versions of these models with different augmentation settings and training data. Beluga used a PANN - cnn14 model (I think it is the same as the original) from his <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183407\" target=\"_blank\">Cornell solution</a>.<br>\nI trained a very similar architecture with different backbones and I used attention head from SED:</li>\n</ul>\n<pre><code>x = ...generate mel-specotrogram...\nx = self.backbone.forward_features(x)\nx = torch.mean(x, dim=2)\nx1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\nx2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\nx = x1 + x2\nx = F.dropout(x, p=0.5, training=self.training)\nx = x.transpose(1, 2)\nx = F.relu_(self.fc1(x))\nx = x.transpose(1, 2)\nx = F.dropout(x, p=0.5, training=self.training)\n(clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\nsegmentwise_output = segmentwise_output.transpose(1, 2)\nframewise_output = interpolate(segmentwise_output, self.interpolate_ratio)\noutput_dict = {\n   \"framewise_output\": framewise_output,\n   \"clipwise_output\": clipwise_output,\n}\n</code></pre>\n<h2>Training</h2>\n<p>Nothing special. Our training method was the same for all of the models:</p>\n<ul>\n<li>4 folds</li>\n<li>10 epochs (15 with a higher probability of mixup)</li>\n<li>Adam (1e-3; *0.9 after 5 epochs)</li>\n<li>BCE loss (PANN version)</li>\n<li>We used the weights of the best validation LWRAP epoch for inference.</li>\n</ul>\n<h2>Pseudo labeling</h2>\n<p>After we had an excellent ensembled score on the public LB (0.950), we started to add pseudo labels to our dataset. The result after we re-trained everything was 0.96x.</p>\n<h2>Final Ensemble</h2>\n<p>Our final ensemble had 80+ models (all of them trained with 4-folds) </p>",
      "rawMarkdown": "Thanks, Kaggle and RFCx, for this audio competition, and special thanks to my teammate @gaborfodor Without him, I probably would have given up a long time ago, somewhere at 0.8xx.\n\n## Data preparation\nResampling everything to 32kHz and split the audio files into 3 seconds duration chunks. We used a sliding window with a 1-second step.\n\n## Collecting more label\nThe key to our result was that Beluga collected tons of training samples manually. He created an awesome annotation application; you can find the details [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220305). *Source code included*.\n\n\nAfter the first batch of manually labeled examples, we quickly achieved 0.93x with an ensemble of a varying number of PANN (cnn14) models. \n\n## Input\nWe used mel-spectrograms as inputs with various `n_bin` (128, 192, 256, 288). Beluga trained PANN-cnn14 models with one input channel. For the other backbones (effnets, resnets, etc) I used three input channels with a simple trick:\n\n- I used different `n_mel` and `n_fft` settings for every channel. E.g. n_mel=(128, 192, 256), n_fft=(1024, 1594, 2048). This results in different height images, so resizing to the same value is necessary.\n\nWe both used `torchlibrosa` to generate the mel-spectrograms.\n\n## Augmentation\nWe used three simple augmentations with different probability:\n\n##### Roll\n```\nnp.roll(y, shift=np.random.randint(0, len(y)))\n```\n\n##### Audio mixup\n```\nw = np.random.uniform(0.3, 0.7)\nmixed = (audio_chunk + rnd_audio_chunk * w) / (1 + w)\nlabel = (label + rnd_label).clip(0, 1)\n```\n\n##### Spec augment\n```\nSpecAugmentation(time_drop_width=16, time_stripes_num=2, freq_drop_width=16, freq_stripes_num=2)\n```\n\n## Architectures\n- PANN - cnn14\n- EfficientNet B0, B1, B2\n- Densenet 121\n- Resnet-50\n- Resnest-50\n- Mobilnet v3 large 100\n\nWe trained many versions of these models with different augmentation settings and training data. Beluga used a PANN - cnn14 model (I think it is the same as the original) from his [Cornell solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183407).\n\nI trained a very similar architecture with different backbones and I used attention head from SED:\n\n```\nx = ...generate mel-specotrogram...\nx = self.backbone.forward_features(x)\n\nx = torch.mean(x, dim=2)\n\nx1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\nx2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\nx = x1 + x2\n\nx = F.dropout(x, p=0.5, training=self.training)\nx = x.transpose(1, 2)\nx = F.relu_(self.fc1(x))\nx = x.transpose(1, 2)\nx = F.dropout(x, p=0.5, training=self.training)\n\n(clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\n\nsegmentwise_output = segmentwise_output.transpose(1, 2)\nframewise_output = interpolate(segmentwise_output, self.interpolate_ratio)\n\noutput_dict = {\n    \"framewise_output\": framewise_output,\n    \"clipwise_output\": clipwise_output,\n}\n```\n\n## Training\nNothing special. Our training method was the same for all of the models:\n- 4 folds\n- 10 epochs (15 with a higher probability of mixup)\n- Adam (1e-3; *0.9 after 5 epochs)\n- BCE loss (PANN version)\n- We used the weights of the best validation LWRAP epoch for inference.\n\n## Pseudo labeling\nAfter we had an excellent ensembled score on the public LB (0.950), we started to add pseudo labels to our dataset. The result after we re-trained everything was 0.96x.\n\n\n## Final Ensemble\nOur final ensemble had 80+ models (all of them trained with 4-folds)",
      "votes": null
    },
    {
      "id": "1208593",
      "postDate": "02/18/2021 10:55:06",
      "content": "<p>Thanks for the writeup, very strong PANN models and augmentation </p>",
      "rawMarkdown": "Thanks for the writeup, very strong PANN models and augmentation",
      "votes": null
    },
    {
      "id": "1208595",
      "postDate": "02/18/2021 10:55:31",
      "content": "<h2>Post Processing</h2>\n<p>For post processing we used  both mean and max predictions to further increase the probability of birds/frogs with frequent calls.</p>\n<pre><code># For each recording\nthresholded_predictions = predictions.copy()\nthresholded_predictions[thresholded_predictions &lt; 0.5] = 0. \n0.7 * predictions.max() + 0.3 * thresholded_predictions.mean()\n</code></pre>",
      "rawMarkdown": "## Post Processing\n\nFor post processing we used  both mean and max predictions to further increase the probability of birds/frogs with frequent calls.\n\n```\n# For each recording\nthresholded_predictions = predictions.copy()\nthresholded_predictions[thresholded_predictions < 0.5] = 0. \n0.7 * predictions.max() + 0.3 * thresholded_predictions.mean()\n```",
      "votes": null
    },
    {
      "id": "1208623",
      "postDate": "02/18/2021 11:27:02",
      "content": "<p>Ho, an yu say more about pseudo labeling.  I could not get an upside out of it.  Maybe it is because when I tried my best model had a public LB score of 0.939, lower than your 0.950?</p>\n<p>Congrats on your final score, and kudos to Gabor for hand labeling data well enough to boost your solution.</p>",
      "rawMarkdown": "Ho, an yu say more about pseudo labeling.  I could not get an upside out of it.  Maybe it is because when I tried my best model had a public LB score of 0.939, lower than your 0.950?\n\nCongrats on your final score, and kudos to Gabor for hand labeling data well enough to boost your solution.",
      "votes": null
    },
    {
      "id": "1208624",
      "postDate": "02/18/2021 11:27:46",
      "content": "<p>This is interesting.  What boost did you get from it?</p>",
      "rawMarkdown": "This is interesting.  What boost did you get from it?",
      "votes": null
    },
    {
      "id": "1208638",
      "postDate": "02/18/2021 11:49:57",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> We predicted both the train and the test set using our best (non-pseudo) ensemble.<br>\nWe had predictions (sigmoid) for 3 seconds chunks using a 1-sec sliding window. For all of the labels, of course.<br>\nI used 0.7 as a threshold for selecting positive chunks. For that chunk, only the birds with 0.7+ prediction considered as positive.<br>\nIt is kind of weak-pseudo labeling.</p>\n<p>My results for one model (Efficientnet B0 - 4 folds avg)<br>\nWith manual labels, I got 0.933 (0.939 private)<br>\nWith manual labels + train only pseudo: 0.949 (0.961 private)<br>\nWith manual labels + test only pseudo: 0.950 (0.956 private)</p>\n<p>The manual+train was better, but it was 4x slower, so I chose the manual + test-only version.</p>\n<p><a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> used different thresholds/selection.</p>",
      "rawMarkdown": "cpmpml We predicted both the train and the test set using our best (non-pseudo) ensemble.\nWe had predictions (sigmoid) for 3 seconds chunks using a 1-sec sliding window. For all of the labels, of course.\nI used 0.7 as a threshold for selecting positive chunks. For that chunk, only the birds with 0.7+ prediction considered as positive.\nIt is kind of weak-pseudo labeling.\n\nMy results for one model (Efficientnet B0 - 4 folds avg)\nWith manual labels, I got 0.933 (0.939 private)\nWith manual labels + train only pseudo: 0.949 (0.961 private)\nWith manual labels + test only pseudo: 0.950 (0.956 private)\n\nThe manual+train was better, but it was 4x slower, so I chose the manual + test-only version.\n\n@gaborfodor used different thresholds/selection.",
      "votes": null
    },
    {
      "id": "1208656",
      "postDate": "02/18/2021 12:07:30",
      "content": "<p>Congratulations🎉</p>",
      "rawMarkdown": "Congratulations🎉",
      "votes": null
    },
    {
      "id": "1208664",
      "postDate": "02/18/2021 12:08:55",
      "content": "<p>It helped ~ 0.003 in the .94 range. After using pseudo labels the boosting effect was negligible</p>",
      "rawMarkdown": "It helped ~ 0.003 in the .94 range. After using pseudo labels the boosting effect was negligible",
      "votes": null
    },
    {
      "id": "1209269",
      "postDate": "02/18/2021 19:37:14",
      "content": "<p>Congratz !<br>\nKudos for having the patience to label extra samples, that's something I'll probably never be capable of doing ^^</p>",
      "rawMarkdown": "Congratz !\nKudos for having the patience to label extra samples, that's something I'll probably never be capable of doing ^^",
      "votes": null
    },
    {
      "id": "1210262",
      "postDate": "02/19/2021 09:55:09",
      "content": "<p>I read somewhere that everyone wants to do ML but no one wants to label data :)</p>\n<p>I would rather collect more data than <strong>only</strong> rely on LB feedback without local validation</p>",
      "rawMarkdown": "I read somewhere that everyone wants to do ML but no one wants to label data :)\n\nI would rather collect more data than **only** rely on LB feedback without local validation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208593,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 10:55:06",
      "content": "<p>Thanks for the writeup, very strong PANN models and augmentation </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208595,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "02/18/2021 10:55:31",
      "content": "<h2>Post Processing</h2>\n<p>For post processing we used  both mean and max predictions to further increase the probability of birds/frogs with frequent calls.</p>\n<pre><code># For each recording\nthresholded_predictions = predictions.copy()\nthresholded_predictions[thresholded_predictions &lt; 0.5] = 0. \n0.7 * predictions.max() + 0.3 * thresholded_predictions.mean()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1208624,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/18/2021 11:27:46",
          "content": "<p>This is interesting.  What boost did you get from it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208664,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/18/2021 12:08:55",
          "content": "<p>It helped ~ 0.003 in the .94 range. After using pseudo labels the boosting effect was negligible</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208623,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 11:27:02",
      "content": "<p>Ho, an yu say more about pseudo labeling.  I could not get an upside out of it.  Maybe it is because when I tried my best model had a public LB score of 0.939, lower than your 0.950?</p>\n<p>Congrats on your final score, and kudos to Gabor for hand labeling data well enough to boost your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208638,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/18/2021 11:49:57",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> We predicted both the train and the test set using our best (non-pseudo) ensemble.<br>\nWe had predictions (sigmoid) for 3 seconds chunks using a 1-sec sliding window. For all of the labels, of course.<br>\nI used 0.7 as a threshold for selecting positive chunks. For that chunk, only the birds with 0.7+ prediction considered as positive.<br>\nIt is kind of weak-pseudo labeling.</p>\n<p>My results for one model (Efficientnet B0 - 4 folds avg)<br>\nWith manual labels, I got 0.933 (0.939 private)<br>\nWith manual labels + train only pseudo: 0.949 (0.961 private)<br>\nWith manual labels + test only pseudo: 0.950 (0.956 private)</p>\n<p>The manual+train was better, but it was 4x slower, so I chose the manual + test-only version.</p>\n<p><a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> used different thresholds/selection.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208656,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "02/18/2021 12:07:30",
      "content": "<p>Congratulations🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209269,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 19:37:14",
      "content": "<p>Congratz !<br>\nKudos for having the patience to label extra samples, that's something I'll probably never be capable of doing ^^</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210262,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/19/2021 09:55:09",
          "content": "<p>I read somewhere that everyone wants to do ML but no one wants to label data :)</p>\n<p>I would rather collect more data than <strong>only</strong> rely on LB feedback without local validation</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1208483": "Thanks, Kaggle and RFCx, for this audio competition, and special thanks to my teammate @gaborfodor Without him, I probably would have given up a long time ago, somewhere at 0.8xx.\n\n## Data preparation\nResampling everything to 32kHz and split the audio files into 3 seconds duration chunks. We used a sliding window with a 1-second step.\n\n## Collecting more label\nThe key to our result was that Beluga collected tons of training samples manually. He created an awesome annotation application; you can find the details [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220305). *Source code included*.\n\n\nAfter the first batch of manually labeled examples, we quickly achieved 0.93x with an ensemble of a varying number of PANN (cnn14) models. \n\n## Input\nWe used mel-spectrograms as inputs with various `n_bin` (128, 192, 256, 288). Beluga trained PANN-cnn14 models with one input channel. For the other backbones (effnets, resnets, etc) I used three input channels with a simple trick:\n\n- I used different `n_mel` and `n_fft` settings for every channel. E.g. n_mel=(128, 192, 256), n_fft=(1024, 1594, 2048). This results in different height images, so resizing to the same value is necessary.\n\nWe both used `torchlibrosa` to generate the mel-spectrograms.\n\n## Augmentation\nWe used three simple augmentations with different probability:\n\n##### Roll\n```\nnp.roll(y, shift=np.random.randint(0, len(y)))\n```\n\n##### Audio mixup\n```\nw = np.random.uniform(0.3, 0.7)\nmixed = (audio_chunk + rnd_audio_chunk * w) / (1 + w)\nlabel = (label + rnd_label).clip(0, 1)\n```\n\n##### Spec augment\n```\nSpecAugmentation(time_drop_width=16, time_stripes_num=2, freq_drop_width=16, freq_stripes_num=2)\n```\n\n## Architectures\n- PANN - cnn14\n- EfficientNet B0, B1, B2\n- Densenet 121\n- Resnet-50\n- Resnest-50\n- Mobilnet v3 large 100\n\nWe trained many versions of these models with different augmentation settings and training data. Beluga used a PANN - cnn14 model (I think it is the same as the original) from his [Cornell solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183407).\n\nI trained a very similar architecture with different backbones and I used attention head from SED:\n\n```\nx = ...generate mel-specotrogram...\nx = self.backbone.forward_features(x)\n\nx = torch.mean(x, dim=2)\n\nx1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\nx2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\nx = x1 + x2\n\nx = F.dropout(x, p=0.5, training=self.training)\nx = x.transpose(1, 2)\nx = F.relu_(self.fc1(x))\nx = x.transpose(1, 2)\nx = F.dropout(x, p=0.5, training=self.training)\n\n(clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\n\nsegmentwise_output = segmentwise_output.transpose(1, 2)\nframewise_output = interpolate(segmentwise_output, self.interpolate_ratio)\n\noutput_dict = {\n    \"framewise_output\": framewise_output,\n    \"clipwise_output\": clipwise_output,\n}\n```\n\n## Training\nNothing special. Our training method was the same for all of the models:\n- 4 folds\n- 10 epochs (15 with a higher probability of mixup)\n- Adam (1e-3; *0.9 after 5 epochs)\n- BCE loss (PANN version)\n- We used the weights of the best validation LWRAP epoch for inference.\n\n## Pseudo labeling\nAfter we had an excellent ensembled score on the public LB (0.950), we started to add pseudo labels to our dataset. The result after we re-trained everything was 0.96x.\n\n\n## Final Ensemble\nOur final ensemble had 80+ models (all of them trained with 4-folds)",
    "1208593": "Thanks for the writeup, very strong PANN models and augmentation",
    "1208595": "## Post Processing\n\nFor post processing we used  both mean and max predictions to further increase the probability of birds/frogs with frequent calls.\n\n```\n# For each recording\nthresholded_predictions = predictions.copy()\nthresholded_predictions[thresholded_predictions < 0.5] = 0. \n0.7 * predictions.max() + 0.3 * thresholded_predictions.mean()\n```",
    "1208623": "Ho, an yu say more about pseudo labeling.  I could not get an upside out of it.  Maybe it is because when I tried my best model had a public LB score of 0.939, lower than your 0.950?\n\nCongrats on your final score, and kudos to Gabor for hand labeling data well enough to boost your solution.",
    "1208624": "This is interesting.  What boost did you get from it?",
    "1208638": "cpmpml We predicted both the train and the test set using our best (non-pseudo) ensemble.\nWe had predictions (sigmoid) for 3 seconds chunks using a 1-sec sliding window. For all of the labels, of course.\nI used 0.7 as a threshold for selecting positive chunks. For that chunk, only the birds with 0.7+ prediction considered as positive.\nIt is kind of weak-pseudo labeling.\n\nMy results for one model (Efficientnet B0 - 4 folds avg)\nWith manual labels, I got 0.933 (0.939 private)\nWith manual labels + train only pseudo: 0.949 (0.961 private)\nWith manual labels + test only pseudo: 0.950 (0.956 private)\n\nThe manual+train was better, but it was 4x slower, so I chose the manual + test-only version.\n\n@gaborfodor used different thresholds/selection.",
    "1208656": "Congratulations🎉",
    "1208664": "It helped ~ 0.003 in the .94 range. After using pseudo labels the boosting effect was negligible",
    "1209269": "Congratz !\nKudos for having the patience to label extra samples, that's something I'll probably never be capable of doing ^^",
    "1210262": "I read somewhere that everyone wants to do ML but no one wants to label data :)\n\nI would rather collect more data than **only** rely on LB feedback without local validation"
  },
  "source": "meta"
}