{
  "id": 243324,
  "title": "9th Place solution",
  "url": "/competitions/birdclef-2021/writeups/tezdhar-9th-place-solution",
  "author_name": "",
  "post_date": "2021-06-02T05:30:26.403Z",
  "votes": 43,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks to all the organizers for hosting such a challenging competition. Also, thanks to all other competitors for sharing valuable information in forums and notebooks. Special thanks to <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> for sharing his code on github from last competition.</p>\n<p>Prayers for all the folks and their families suffering from Covid. </p>\n<h3>Solution Summary:</h3>\n<p>For most of my models, I used 5-7 second clips for training. I gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions. </p>\n<h4>Ensembling</h4>\n<p>Mostly, the strategy was to get as much diversity possible into the ensemble to make predictions more robust.</p>\n<ul>\n<li>Melspecs with different temporal resolution (hop_length - 200 and 320)</li>\n<li>resnest50, efficientnetb0 and densenet121 backbones</li>\n<li>Noise augmentation (white noise, pink noise, band noise, nocall clips) same as what <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> used in previous competition</li>\n</ul>\n<h4>Post processing</h4>\n<p>This was the trickiest part. We had following information at hand in addition to raw probabilities for each 5s clip:</p>\n<ul>\n<li>Maximum probability of bird over whole audio clip</li>\n<li>Average probability of bird over whole audio clip</li>\n<li>Count of birds from <code>train_metadata.csv</code> within +/- x deg of soundscape location</li>\n<li>Maximum probability of bird over whole day</li>\n</ul>\n<p>Initially I tried fitting a meta estimator (LightGBM and RandomForest, LogisticRegression) over above variable but results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).<br>\nCustom rule was something like: <code>p1 * a1 + np.clip(p2, 0, t1) * a2 + ...</code></p>\n<h4>Other tricks</h4>\n<ul>\n<li>Raise melspecs for different models to different powers to add little more diversity</li>\n<li>Squeeze width of test soundscapes by 2-5% (mostly to reverse far field effects)</li>\n</ul>\n<h4>Things that didn't work</h4>\n<ul>\n<li>Using last name from common name (e.g. <code>parakeet</code> from <code>Orange billed parakeet</code>) as an additional label during training.</li>\n<li>Different Attentions on backbone feature map (I guess with just 5-7 s clips, adding attention confuses the model even more instead of helping. Simple sum/mean over seems to work best.)</li>\n<li>Bigger models, I had really hard time getting effcientnet-b2/b4 to converge. I just gave up on them due to lack on time and resources.</li>\n<li>Manual labelling and class balancing. Looking at results from my models, it was quite evident that birds with less no of samples were not getting predicted (e.g. <code>hofwoo1</code>). So, I split them into 10s clips  and hand labelled them to increase their sample size but this didnt help. Also, class balancing gave lower validation loss to ditched it as well.</li>\n</ul>",
  "messages": [
    {
      "id": "1332331",
      "postDate": "06/02/2021 04:32:32",
      "content": "<p>Thanks to all the organizers for hosting such a challenging competition. Also, thanks to all other competitors for sharing valuable information in forums and notebooks. Special thanks to <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> for sharing his code on github from last competition.</p>\n<p>Prayers for all the folks and their families suffering from Covid. </p>\n<h3>Solution Summary:</h3>\n<p>For most of my models, I used 5-7 second clips for training. I gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions. </p>\n<h4>Ensembling</h4>\n<p>Mostly, the strategy was to get as much diversity possible into the ensemble to make predictions more robust.</p>\n<ul>\n<li>Melspecs with different temporal resolution (hop_length - 200 and 320)</li>\n<li>resnest50, efficientnetb0 and densenet121 backbones</li>\n<li>Noise augmentation (white noise, pink noise, band noise, nocall clips) same as what <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> used in previous competition</li>\n</ul>\n<h4>Post processing</h4>\n<p>This was the trickiest part. We had following information at hand in addition to raw probabilities for each 5s clip:</p>\n<ul>\n<li>Maximum probability of bird over whole audio clip</li>\n<li>Average probability of bird over whole audio clip</li>\n<li>Count of birds from <code>train_metadata.csv</code> within +/- x deg of soundscape location</li>\n<li>Maximum probability of bird over whole day</li>\n</ul>\n<p>Initially I tried fitting a meta estimator (LightGBM and RandomForest, LogisticRegression) over above variable but results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).<br>\nCustom rule was something like: <code>p1 * a1 + np.clip(p2, 0, t1) * a2 + ...</code></p>\n<h4>Other tricks</h4>\n<ul>\n<li>Raise melspecs for different models to different powers to add little more diversity</li>\n<li>Squeeze width of test soundscapes by 2-5% (mostly to reverse far field effects)</li>\n</ul>\n<h4>Things that didn't work</h4>\n<ul>\n<li>Using last name from common name (e.g. <code>parakeet</code> from <code>Orange billed parakeet</code>) as an additional label during training.</li>\n<li>Different Attentions on backbone feature map (I guess with just 5-7 s clips, adding attention confuses the model even more instead of helping. Simple sum/mean over seems to work best.)</li>\n<li>Bigger models, I had really hard time getting effcientnet-b2/b4 to converge. I just gave up on them due to lack on time and resources.</li>\n<li>Manual labelling and class balancing. Looking at results from my models, it was quite evident that birds with less no of samples were not getting predicted (e.g. <code>hofwoo1</code>). So, I split them into 10s clips  and hand labelled them to increase their sample size but this didnt help. Also, class balancing gave lower validation loss to ditched it as well.</li>\n</ul>",
      "rawMarkdown": "Thanks to all the organizers for hosting such a challenging competition. Also, thanks to all other competitors for sharing valuable information in forums and notebooks. Special thanks to @vlomme for sharing his code on github from last competition.\n\nPrayers for all the folks and their families suffering from Covid. \n\n### Solution Summary:\n\nFor most of my models, I used 5-7 second clips for training. I gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions. \n\n#### Ensembling\nMostly, the strategy was to get as much diversity possible into the ensemble to make predictions more robust.\n\n* Melspecs with different temporal resolution (hop_length - 200 and 320)\n* resnest50, efficientnetb0 and densenet121 backbones\n* Noise augmentation (white noise, pink noise, band noise, nocall clips) same as what @vlomme used in previous competition\n\n#### Post processing\n\nThis was the trickiest part. We had following information at hand in addition to raw probabilities for each 5s clip:\n* Maximum probability of bird over whole audio clip\n* Average probability of bird over whole audio clip\n* Count of birds from `train_metadata.csv` within +/- x deg of soundscape location\n* Maximum probability of bird over whole day\n\nInitially I tried fitting a meta estimator (LightGBM and RandomForest, LogisticRegression) over above variable but results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).\nCustom rule was something like: `p1 * a1 + np.clip(p2, 0, t1) * a2 + ...`\n\n#### Other tricks\n* Raise melspecs for different models to different powers to add little more diversity\n* Squeeze width of test soundscapes by 2-5% (mostly to reverse far field effects)\n\n#### Things that didn't work\n* Using last name from common name (e.g. `parakeet` from `Orange billed parakeet`) as an additional label during training.\n* Different Attentions on backbone feature map (I guess with just 5-7 s clips, adding attention confuses the model even more instead of helping. Simple sum/mean over seems to work best.)\n* Bigger models, I had really hard time getting effcientnet-b2/b4 to converge. I just gave up on them due to lack on time and resources.\n* Manual labelling and class balancing. Looking at results from my models, it was quite evident that birds with less no of samples were not getting predicted (e.g. `hofwoo1`). So, I split them into 10s clips  and hand labelled them to increase their sample size but this didnt help. Also, class balancing gave lower validation loss to ditched it as well.",
      "votes": null
    },
    {
      "id": "1332422",
      "postDate": "06/02/2021 05:50:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> congrats on the solo gold , thanks for sharing the solution , I had two questions , it would be great if you can answer them :</p>\n<ul>\n<li>What was your training strategy for efficientnet did you use a custom head , etc because we tried a lot but it didn't work with effnet</li>\n<li>Could you elaborate a bit more on post processing </li>\n</ul>",
      "rawMarkdown": "Hi @tezdhar congrats on the solo gold , thanks for sharing the solution , I had two questions , it would be great if you can answer them :\n* What was your training strategy for efficientnet did you use a custom head , etc because we tried a lot but it didn't work with effnet\n* Could you elaborate a bit more on post processing",
      "votes": null
    },
    {
      "id": "1332428",
      "postDate": "06/02/2021 05:55:21",
      "content": "<p>It's nice that my decision was useful to someone. And I'm even more pleased to see that it was the basis for the first place decision in the public rankings.</p>",
      "rawMarkdown": "It's nice that my decision was useful to someone. And I'm even more pleased to see that it was the basis for the first place decision in the public rankings.",
      "votes": null
    },
    {
      "id": "1332569",
      "postDate": "06/02/2021 07:34:48",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> !</p>\n<ul>\n<li>Discriminative LR helps efficientnet converge. Backbone lr for efficientnets were 0.1 or 0.5 times max_lr. Works with different heads. </li>\n<li>Post processing function \\</li>\n</ul>\n<pre><code>def post_process2(probas, test, bird_loc_cnts, max_w=0.5, mean_w=0.5, cnt_w=1, max_w2=0.1):\n    proba = probas.copy()\n    dates = test.filepaths.str.split('/').str.get(-1).str.slice(-12,-4).values\n    audio_ids = test.audio_id.values\n    proba_date_max = groupby_np(dates, probas, np.max)\n    proba_audio_max = groupby_np(audio_ids, probas, np.max)\n    proba_audio_mean = groupby_np(audio_ids, probas, np.mean)\n\n    for audio_id in test.audio_id.unique():\n        rows = test.audio_id.values == audio_id\n        site = test.loc[rows, 'site'].iloc[0]\n        bird_cnts = bird_loc_cnts[site]\n        cnt_prob1 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 1) for i in range(398)]).reshape(1, -1), sum(rows), 0))\n        cnt_prob2 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 100) for i in range(398)]).reshape(1, -1), sum(rows), 0))/100\n        # max_probs = np.max(proba[rows], axis=0, keepdims=True)\n        # thresh_cnt = np.sum(proba[rows] &gt; thresh, axis=0, keepdims=True)\n        # mean_probs = np.mean(proba[rows], axis=0, keepdims=True)\n        max_probs = proba_audio_max[rows]\n        mean_probs = proba_audio_mean[rows]\n        max_probs2 = proba_date_max[rows]\n        if cnt_w &gt; 0:\n            proba[rows] *= cnt_prob1*1.0\n        proba[rows] = proba[rows] + max_w*np.clip(max_probs, 0, 0.67) + max_w2*np.clip(max_probs2, 0, 0.67) + mean_w*np.clip(mean_probs, 0, 0.1) + cnt_w*np.clip(cnt_prob2, 0, 0.01) # + thresh_w*np.clip(thresh_cnt, 0, 5)/5.0\n    proba[:, 397] = probas[:, 397]\n    return proba\n</code></pre>",
      "rawMarkdown": "Thanks @tanulsingh077 !\n\n* Discriminative LR helps efficientnet converge. Backbone lr for efficientnets were 0.1 or 0.5 times max_lr. Works with different heads. \n* Post processing function \\\n```\ndef post_process2(probas, test, bird_loc_cnts, max_w=0.5, mean_w=0.5, cnt_w=1, max_w2=0.1):\n    proba = probas.copy()\n    dates = test.filepaths.str.split('/').str.get(-1).str.slice(-12,-4).values\n    audio_ids = test.audio_id.values\n    proba_date_max = groupby_np(dates, probas, np.max)\n    proba_audio_max = groupby_np(audio_ids, probas, np.max)\n    proba_audio_mean = groupby_np(audio_ids, probas, np.mean)\n\n    for audio_id in test.audio_id.unique():\n        rows = test.audio_id.values == audio_id\n        site = test.loc[rows, 'site'].iloc[0]\n        bird_cnts = bird_loc_cnts[site]\n        cnt_prob1 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 1) for i in range(398)]).reshape(1, -1), sum(rows), 0))\n        cnt_prob2 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 100) for i in range(398)]).reshape(1, -1), sum(rows), 0))/100\n        # max_probs = np.max(proba[rows], axis=0, keepdims=True)\n        # thresh_cnt = np.sum(proba[rows] > thresh, axis=0, keepdims=True)\n        # mean_probs = np.mean(proba[rows], axis=0, keepdims=True)\n        max_probs = proba_audio_max[rows]\n        mean_probs = proba_audio_mean[rows]\n        max_probs2 = proba_date_max[rows]\n        if cnt_w > 0:\n            proba[rows] *= cnt_prob1*1.0\n        proba[rows] = proba[rows] + max_w*np.clip(max_probs, 0, 0.67) + max_w2*np.clip(max_probs2, 0, 0.67) + mean_w*np.clip(mean_probs, 0, 0.1) + cnt_w*np.clip(cnt_prob2, 0, 0.01) # + thresh_w*np.clip(thresh_cnt, 0, 5)/5.0\n    proba[:, 397] = probas[:, 397]\n    return proba\n```",
      "votes": null
    },
    {
      "id": "1332613",
      "postDate": "06/02/2021 08:01:22",
      "content": "<p>Congrats on your result.  I was watching you climb the LB and thought you would finish well.</p>\n<blockquote>\n  <p>results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).</p>\n</blockquote>\n<p>I did the same mistake probably.  My post processing helped a lot, maybe too much…</p>",
      "rawMarkdown": "Congrats on your result.  I was watching you climb the LB and thought you would finish well.\n\n> results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).\n\nI did the same mistake probably.  My post processing helped a lot, maybe too much...",
      "votes": null
    },
    {
      "id": "1332617",
      "postDate": "06/02/2021 08:07:35",
      "content": "<p>Congrats on solo gold, great effort!</p>\n<p>I am a bit surprised that both you and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> only worked on 5sec crops if I understand that correctly. It appeared to us that there is way too much noise if you only pick such a short sub-sample.</p>",
      "rawMarkdown": "Congrats on solo gold, great effort!\n\nI am a bit surprised that both you and @cpmpml only worked on 5sec crops if I understand that correctly. It appeared to us that there is way too much noise if you only pick such a short sub-sample.",
      "votes": null
    },
    {
      "id": "1332624",
      "postDate": "06/02/2021 08:11:12",
      "content": "<p>Well, it matches what we have to predict.  But we cannot capture longer clip dynamics and have to rely on post processing for it.  Mine is similar to what is described here.  The trick is that this worked very well on CV and public LB.  Maybe we were fooled by it.</p>\n<p>When I saw how much tuning I had ot do on postprocessing I thought it would be better to have it learned by a model.  This leads to SED like model.  I only started on it the last day, which was hopeless of course.</p>",
      "rawMarkdown": "Well, it matches what we have to predict.  But we cannot capture longer clip dynamics and have to rely on post processing for it.  Mine is similar to what is described here.  The trick is that this worked very well on CV and public LB.  Maybe we were fooled by it.\n\nWhen I saw how much tuning I had ot do on postprocessing I thought it would be better to have it learned by a model.  This leads to SED like model.  I only started on it the last day, which was hopeless of course.",
      "votes": null
    },
    {
      "id": "1332627",
      "postDate": "06/02/2021 08:12:15",
      "content": "<p>But if you crop only 5 seconds you have a lot more wrong labels while training because we only have weak labels. </p>",
      "rawMarkdown": "But if you crop only 5 seconds you have a lot more wrong labels while training because we only have weak labels.",
      "votes": null
    },
    {
      "id": "1332660",
      "postDate": "06/02/2021 08:30:00",
      "content": "<p><a href=\"https://www.kaggle.com/psi\" target=\"_blank\">@psi</a> sorry I am coming in between your discussions , we have train soundscapes which are  30secs-3min long , and I think <a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a> is cropping the start of the audio and the end of the audio from these based on his hypothesis that the train soundscapes were cropped to the start of birdcall start and hence this would reduce noise.</p>\n<p>In last year's comp people used a lot of things to counter the issue mentioned by you , theo used better cropping based on oof's score , vladmir used RMS for better cropping</p>",
      "rawMarkdown": "psi sorry I am coming in between your discussions , we have train soundscapes which are  30secs-3min long , and I think @cpmp is cropping the start of the audio and the end of the audio from these based on his hypothesis that the train soundscapes were cropped to the start of birdcall start and hence this would reduce noise.\n\nIn last year's comp people used a lot of things to counter the issue mentioned by you , theo used better cropping based on oof's score , vladmir used RMS for better cropping",
      "votes": null
    },
    {
      "id": "1332662",
      "postDate": "06/02/2021 08:30:15",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>   As in Cornell I used first 5 seconds and last 5 seconds.  Rationale is given in this post:</p>\n<p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/230733#1265442\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/230733#1265442</a></p>\n<p>I commented I used the same in Cornell competition ;)  And I reused it here as well.</p>",
      "rawMarkdown": "philippsinger   As in Cornell I used first 5 seconds and last 5 seconds.  Rationale is given in this post:\n\nhttps://www.kaggle.com/c/birdclef-2021/discussion/230733#1265442\n\nI commented I used the same in Cornell competition ;)  And I reused it here as well.",
      "votes": null
    },
    {
      "id": "1332680",
      "postDate": "06/02/2021 08:38:59",
      "content": "<p>Makes sense to reduce the noise. Did you still try larger crop lengths? </p>",
      "rawMarkdown": "Makes sense to reduce the noise. Did you still try larger crop lengths?",
      "votes": null
    },
    {
      "id": "1332698",
      "postDate": "06/02/2021 08:48:05",
      "content": "<p>Thanks for providing the function 😊</p>",
      "rawMarkdown": "Thanks for providing the function 😊",
      "votes": null
    },
    {
      "id": "1332707",
      "postDate": "06/02/2021 08:52:31",
      "content": "<p>I looked into many train short audios and noticed in many files the 1st or last seconds are people speaking, some language I don<code>t understand, so we don</code>t use assumption to crop 1st or last 5 seconds. Instead use energy + pseudo based probs cropping.</p>",
      "rawMarkdown": "I looked into many train short audios and noticed in many files the 1st or last seconds are people speaking, some language I don`t understand, so we don`t use assumption to crop 1st or last 5 seconds. Instead use energy + pseudo based probs cropping.",
      "votes": null
    },
    {
      "id": "1332720",
      "postDate": "06/02/2021 09:01:39",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! Yes, I did try training on 30s clips few days before deadline with same noise setup. But that did not give me good results on train soundscapes (~0.68 f1) which was too low to be even included in ensemble. <br>\nIt could be I had some bug in my pipeline or using too much augmentation was causing some issues. Anyway, I guess I decided to drop it too early without much experimentation.</p>",
      "rawMarkdown": "Thanks @philippsinger! Yes, I did try training on 30s clips few days before deadline with same noise setup. But that did not give me good results on train soundscapes (~0.68 f1) which was too low to be even included in ensemble. \nIt could be I had some bug in my pipeline or using too much augmentation was causing some issues. Anyway, I guess I decided to drop it too early without much experimentation.",
      "votes": null
    },
    {
      "id": "1332725",
      "postDate": "06/02/2021 09:05:30",
      "content": "<p>Thank you :). <br>\nWe fell for the mirage created by public LB. Hopefully, I am able to do better in my next competitions</p>",
      "rawMarkdown": "Thank you :). \nWe fell for the mirage created by public LB. Hopefully, I am able to do better in my next competitions",
      "votes": null
    },
    {
      "id": "1332732",
      "postDate": "06/02/2021 09:09:03",
      "content": "<p>One more thing, you could use different thresholds for bird and nocall. For example, you could make prediction like this <code>chswar nocall</code>. This helped slightly.</p>",
      "rawMarkdown": "One more thing, you could use different thresholds for bird and nocall. For example, you could make prediction like this `chswar nocall`. This helped slightly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1332422,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "06/02/2021 05:50:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> congrats on the solo gold , thanks for sharing the solution , I had two questions , it would be great if you can answer them :</p>\n<ul>\n<li>What was your training strategy for efficientnet did you use a custom head , etc because we tried a lot but it didn't work with effnet</li>\n<li>Could you elaborate a bit more on post processing </li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1332569,
          "author_name": "tezdhar",
          "author_url": "",
          "post_date": "06/02/2021 07:34:48",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> !</p>\n<ul>\n<li>Discriminative LR helps efficientnet converge. Backbone lr for efficientnets were 0.1 or 0.5 times max_lr. Works with different heads. </li>\n<li>Post processing function \\</li>\n</ul>\n<pre><code>def post_process2(probas, test, bird_loc_cnts, max_w=0.5, mean_w=0.5, cnt_w=1, max_w2=0.1):\n    proba = probas.copy()\n    dates = test.filepaths.str.split('/').str.get(-1).str.slice(-12,-4).values\n    audio_ids = test.audio_id.values\n    proba_date_max = groupby_np(dates, probas, np.max)\n    proba_audio_max = groupby_np(audio_ids, probas, np.max)\n    proba_audio_mean = groupby_np(audio_ids, probas, np.mean)\n\n    for audio_id in test.audio_id.unique():\n        rows = test.audio_id.values == audio_id\n        site = test.loc[rows, 'site'].iloc[0]\n        bird_cnts = bird_loc_cnts[site]\n        cnt_prob1 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 1) for i in range(398)]).reshape(1, -1), sum(rows), 0))\n        cnt_prob2 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 100) for i in range(398)]).reshape(1, -1), sum(rows), 0))/100\n        # max_probs = np.max(proba[rows], axis=0, keepdims=True)\n        # thresh_cnt = np.sum(proba[rows] &gt; thresh, axis=0, keepdims=True)\n        # mean_probs = np.mean(proba[rows], axis=0, keepdims=True)\n        max_probs = proba_audio_max[rows]\n        mean_probs = proba_audio_mean[rows]\n        max_probs2 = proba_date_max[rows]\n        if cnt_w &gt; 0:\n            proba[rows] *= cnt_prob1*1.0\n        proba[rows] = proba[rows] + max_w*np.clip(max_probs, 0, 0.67) + max_w2*np.clip(max_probs2, 0, 0.67) + mean_w*np.clip(mean_probs, 0, 0.1) + cnt_w*np.clip(cnt_prob2, 0, 0.01) # + thresh_w*np.clip(thresh_cnt, 0, 5)/5.0\n    proba[:, 397] = probas[:, 397]\n    return proba\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332698,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/02/2021 08:48:05",
          "content": "<p>Thanks for providing the function 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332428,
      "author_name": "vlomme",
      "author_url": "",
      "post_date": "06/02/2021 05:55:21",
      "content": "<p>It's nice that my decision was useful to someone. And I'm even more pleased to see that it was the basis for the first place decision in the public rankings.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332613,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/02/2021 08:01:22",
      "content": "<p>Congrats on your result.  I was watching you climb the LB and thought you would finish well.</p>\n<blockquote>\n  <p>results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).</p>\n</blockquote>\n<p>I did the same mistake probably.  My post processing helped a lot, maybe too much…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332725,
          "author_name": "tezdhar",
          "author_url": "",
          "post_date": "06/02/2021 09:05:30",
          "content": "<p>Thank you :). <br>\nWe fell for the mirage created by public LB. Hopefully, I am able to do better in my next competitions</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332617,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "06/02/2021 08:07:35",
      "content": "<p>Congrats on solo gold, great effort!</p>\n<p>I am a bit surprised that both you and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> only worked on 5sec crops if I understand that correctly. It appeared to us that there is way too much noise if you only pick such a short sub-sample.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332624,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/02/2021 08:11:12",
          "content": "<p>Well, it matches what we have to predict.  But we cannot capture longer clip dynamics and have to rely on post processing for it.  Mine is similar to what is described here.  The trick is that this worked very well on CV and public LB.  Maybe we were fooled by it.</p>\n<p>When I saw how much tuning I had ot do on postprocessing I thought it would be better to have it learned by a model.  This leads to SED like model.  I only started on it the last day, which was hopeless of course.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332627,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/02/2021 08:12:15",
          "content": "<p>But if you crop only 5 seconds you have a lot more wrong labels while training because we only have weak labels. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332660,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "06/02/2021 08:30:00",
          "content": "<p><a href=\"https://www.kaggle.com/psi\" target=\"_blank\">@psi</a> sorry I am coming in between your discussions , we have train soundscapes which are  30secs-3min long , and I think <a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a> is cropping the start of the audio and the end of the audio from these based on his hypothesis that the train soundscapes were cropped to the start of birdcall start and hence this would reduce noise.</p>\n<p>In last year's comp people used a lot of things to counter the issue mentioned by you , theo used better cropping based on oof's score , vladmir used RMS for better cropping</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332662,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/02/2021 08:30:15",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>   As in Cornell I used first 5 seconds and last 5 seconds.  Rationale is given in this post:</p>\n<p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/230733#1265442\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/230733#1265442</a></p>\n<p>I commented I used the same in Cornell competition ;)  And I reused it here as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332680,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/02/2021 08:38:59",
          "content": "<p>Makes sense to reduce the noise. Did you still try larger crop lengths? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332707,
          "author_name": "superchenhao",
          "author_url": "",
          "post_date": "06/02/2021 08:52:31",
          "content": "<p>I looked into many train short audios and noticed in many files the 1st or last seconds are people speaking, some language I don<code>t understand, so we don</code>t use assumption to crop 1st or last 5 seconds. Instead use energy + pseudo based probs cropping.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332720,
          "author_name": "tezdhar",
          "author_url": "",
          "post_date": "06/02/2021 09:01:39",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! Yes, I did try training on 30s clips few days before deadline with same noise setup. But that did not give me good results on train soundscapes (~0.68 f1) which was too low to be even included in ensemble. <br>\nIt could be I had some bug in my pipeline or using too much augmentation was causing some issues. Anyway, I guess I decided to drop it too early without much experimentation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332732,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "06/02/2021 09:09:03",
      "content": "<p>One more thing, you could use different thresholds for bird and nocall. For example, you could make prediction like this <code>chswar nocall</code>. This helped slightly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1332331": "Thanks to all the organizers for hosting such a challenging competition. Also, thanks to all other competitors for sharing valuable information in forums and notebooks. Special thanks to @vlomme for sharing his code on github from last competition.\n\nPrayers for all the folks and their families suffering from Covid. \n\n### Solution Summary:\n\nFor most of my models, I used 5-7 second clips for training. I gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions. \n\n#### Ensembling\nMostly, the strategy was to get as much diversity possible into the ensemble to make predictions more robust.\n\n* Melspecs with different temporal resolution (hop_length - 200 and 320)\n* resnest50, efficientnetb0 and densenet121 backbones\n* Noise augmentation (white noise, pink noise, band noise, nocall clips) same as what @vlomme used in previous competition\n\n#### Post processing\n\nThis was the trickiest part. We had following information at hand in addition to raw probabilities for each 5s clip:\n* Maximum probability of bird over whole audio clip\n* Average probability of bird over whole audio clip\n* Count of birds from `train_metadata.csv` within +/- x deg of soundscape location\n* Maximum probability of bird over whole day\n\nInitially I tried fitting a meta estimator (LightGBM and RandomForest, LogisticRegression) over above variable but results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).\nCustom rule was something like: `p1 * a1 + np.clip(p2, 0, t1) * a2 + ...`\n\n#### Other tricks\n* Raise melspecs for different models to different powers to add little more diversity\n* Squeeze width of test soundscapes by 2-5% (mostly to reverse far field effects)\n\n#### Things that didn't work\n* Using last name from common name (e.g. `parakeet` from `Orange billed parakeet`) as an additional label during training.\n* Different Attentions on backbone feature map (I guess with just 5-7 s clips, adding attention confuses the model even more instead of helping. Simple sum/mean over seems to work best.)\n* Bigger models, I had really hard time getting effcientnet-b2/b4 to converge. I just gave up on them due to lack on time and resources.\n* Manual labelling and class balancing. Looking at results from my models, it was quite evident that birds with less no of samples were not getting predicted (e.g. `hofwoo1`). So, I split them into 10s clips  and hand labelled them to increase their sample size but this didnt help. Also, class balancing gave lower validation loss to ditched it as well.",
    "1332422": "Hi @tezdhar congrats on the solo gold , thanks for sharing the solution , I had two questions , it would be great if you can answer them :\n* What was your training strategy for efficientnet did you use a custom head , etc because we tried a lot but it didn't work with effnet\n* Could you elaborate a bit more on post processing",
    "1332428": "It's nice that my decision was useful to someone. And I'm even more pleased to see that it was the basis for the first place decision in the public rankings.",
    "1332569": "Thanks @tanulsingh077 !\n\n* Discriminative LR helps efficientnet converge. Backbone lr for efficientnets were 0.1 or 0.5 times max_lr. Works with different heads. \n* Post processing function \\\n```\ndef post_process2(probas, test, bird_loc_cnts, max_w=0.5, mean_w=0.5, cnt_w=1, max_w2=0.1):\n    proba = probas.copy()\n    dates = test.filepaths.str.split('/').str.get(-1).str.slice(-12,-4).values\n    audio_ids = test.audio_id.values\n    proba_date_max = groupby_np(dates, probas, np.max)\n    proba_audio_max = groupby_np(audio_ids, probas, np.max)\n    proba_audio_mean = groupby_np(audio_ids, probas, np.mean)\n\n    for audio_id in test.audio_id.unique():\n        rows = test.audio_id.values == audio_id\n        site = test.loc[rows, 'site'].iloc[0]\n        bird_cnts = bird_loc_cnts[site]\n        cnt_prob1 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 1) for i in range(398)]).reshape(1, -1), sum(rows), 0))\n        cnt_prob2 = np.log1p(np.repeat(np.array([np.clip(bird_cnts.get(INT2CODE[i], 0), 0.0, 100) for i in range(398)]).reshape(1, -1), sum(rows), 0))/100\n        # max_probs = np.max(proba[rows], axis=0, keepdims=True)\n        # thresh_cnt = np.sum(proba[rows] > thresh, axis=0, keepdims=True)\n        # mean_probs = np.mean(proba[rows], axis=0, keepdims=True)\n        max_probs = proba_audio_max[rows]\n        mean_probs = proba_audio_mean[rows]\n        max_probs2 = proba_date_max[rows]\n        if cnt_w > 0:\n            proba[rows] *= cnt_prob1*1.0\n        proba[rows] = proba[rows] + max_w*np.clip(max_probs, 0, 0.67) + max_w2*np.clip(max_probs2, 0, 0.67) + mean_w*np.clip(mean_probs, 0, 0.1) + cnt_w*np.clip(cnt_prob2, 0, 0.01) # + thresh_w*np.clip(thresh_cnt, 0, 5)/5.0\n    proba[:, 397] = probas[:, 397]\n    return proba\n```",
    "1332613": "Congrats on your result.  I was watching you climb the LB and thought you would finish well.\n\n> results using custom rule were better on both validation and public LB (Thinking retrospectively, I think this is where I overfit).\n\nI did the same mistake probably.  My post processing helped a lot, maybe too much...",
    "1332617": "Congrats on solo gold, great effort!\n\nI am a bit surprised that both you and @cpmpml only worked on 5sec crops if I understand that correctly. It appeared to us that there is way too much noise if you only pick such a short sub-sample.",
    "1332624": "Well, it matches what we have to predict.  But we cannot capture longer clip dynamics and have to rely on post processing for it.  Mine is similar to what is described here.  The trick is that this worked very well on CV and public LB.  Maybe we were fooled by it.\n\nWhen I saw how much tuning I had ot do on postprocessing I thought it would be better to have it learned by a model.  This leads to SED like model.  I only started on it the last day, which was hopeless of course.",
    "1332627": "But if you crop only 5 seconds you have a lot more wrong labels while training because we only have weak labels.",
    "1332660": "psi sorry I am coming in between your discussions , we have train soundscapes which are  30secs-3min long , and I think @cpmp is cropping the start of the audio and the end of the audio from these based on his hypothesis that the train soundscapes were cropped to the start of birdcall start and hence this would reduce noise.\n\nIn last year's comp people used a lot of things to counter the issue mentioned by you , theo used better cropping based on oof's score , vladmir used RMS for better cropping",
    "1332662": "philippsinger   As in Cornell I used first 5 seconds and last 5 seconds.  Rationale is given in this post:\n\nhttps://www.kaggle.com/c/birdclef-2021/discussion/230733#1265442\n\nI commented I used the same in Cornell competition ;)  And I reused it here as well.",
    "1332680": "Makes sense to reduce the noise. Did you still try larger crop lengths?",
    "1332698": "Thanks for providing the function 😊",
    "1332707": "I looked into many train short audios and noticed in many files the 1st or last seconds are people speaking, some language I don`t understand, so we don`t use assumption to crop 1st or last 5 seconds. Instead use energy + pseudo based probs cropping.",
    "1332720": "Thanks @philippsinger! Yes, I did try training on 30s clips few days before deadline with same noise setup. But that did not give me good results on train soundscapes (~0.68 f1) which was too low to be even included in ensemble. \nIt could be I had some bug in my pipeline or using too much augmentation was causing some issues. Anyway, I guess I decided to drop it too early without much experimentation.",
    "1332725": "Thank you :). \nWe fell for the mirage created by public LB. Hopefully, I am able to do better in my next competitions",
    "1332732": "One more thing, you could use different thresholds for bird and nocall. For example, you could make prediction like this `chswar nocall`. This helped slightly."
  },
  "source": "meta"
}