{
  "id": 183209,
  "title": "21st Place Write-up [Stack-up]",
  "url": "/competitions/birdsong-recognition/writeups/hope-21st-place-write-up-stack-up",
  "author_name": "",
  "post_date": "2021-09-08T15:30:23.323Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all, I would like to thank kaggle and the organizers for hosting such an interesting competition.</p>\n<p>thanks for my wonderful team <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> and <a href=\"https://www.kaggle.com/tarique7\" target=\"_blank\">@tarique7</a>. this was great learning and collaboration.</p>\n<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> and I worked hard for this from the last 20days</p>\n<h4>[Big Shake-up]</h4>\n<ul>\n<li>LB : 0.533 ---&gt; Private LB : 0.632</li>\n<li>From 364th place ---&gt; 21st place</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2058044%2F0725e3d4482b444098f50d6df08441ef%2FScreenshot%20from%202020-09-16%2006-20-45.png?generation=1600217490267914&amp;alt=media\" alt=\"\"></p>\n<h3>[Summary]</h3>\n<ul>\n<li>converted mp3 audio file into wav with sample_rate 16k, because of colab drive space and birds sing up to 12kHz</li>\n<li>we trained <strong>stage1</strong> model based on 5sec random clip.</li>\n<li>It gives average of 0.65+ AMP on 5-fold</li>\n<li>we created 1sec audio clips dataset.</li>\n<li>predicted 5fold <strong>stage1</strong> models on 1sec dataset.</li>\n<li>selected <code>df[(df.ebird_code == df.pred_code) &amp; (df.pred_prob &gt;= 0.5)]</code> selected those 1sec clips for <strong>stage3</strong></li>\n<li>we trained <strong>stage2</strong> model on topof <strong>stage1</strong> using public data</li>\n<li>again predicted <strong>stage1</strong> and <strong>stage2</strong> models on public 1sec clips</li>\n<li>again select <code>df2[(df2.ebird_code == df2.pred_code) &amp; (df2.pred_prob &gt;= 0.5)]</code>selected those 1sec clips for <strong>stage3</strong></li>\n<li><strong>stage3</strong> dataset is <code>stage3_df = df.append(df2)</code> we endup with 612K 1sec clips with sudo labels</li>\n<li>we created 1sec noise clips using PANN <strong>Cnn14_16k</strong> model.</li>\n<li>we predict <strong>Cnn14_16k</strong> model on 1sec dataset and select some noise labels from PANN labels </li>\n<li>those labels are <code>['Silence', 'White noise', 'Vachical', 'Speech', 'Pink noise', 'Tick-tock', 'Wind noise (microphone)','Stream','Raindrop','Wind','Rain', ...... ]</code> selected those labels as noise labels</li>\n<li>based on those noise labels and <strong>stage1</strong> model predicted probabilities we select 1sec noise clips data.</li>\n<li><code>noise_df = df[(df.filename.isin(noise_labels_df.filename)) &amp; df.pred_prob &lt; 0.4]</code></li>\n<li>we end up with 103k noise 1sec clips you find dataset <a href=\"https://www.kaggle.com/gopidurgaprasad/birdsong-stage1-1sec-sudo-noise\" target=\"_blank\">link</a></li>\n<li>we know in this competition our main goal is to predict a mix of bird calls.</li>\n<li>now the main part, at the end we have <strong>stage1</strong>, <strong>stage2</strong> models, <strong>1sec</strong> dataset with Sudo labels, and <strong>1sec noise data</strong>. we trained the <strong>stage3</strong> model using all of those.</li>\n<li>now in front of us, we need to build <strong>CV</strong> and train a model that more reliable on predicting a mix of bird calls.</li>\n</ul>\n<h3>[CV]</h3>\n<ul>\n<li>we created a <strong>cv</strong> based on <strong>1sec bird calls</strong> and <strong>1sec noise data</strong></li>\n<li>In the end, we need to predict for <strong>5sec</strong> so we take 5 random birdcalls and noise stack them and give labels based on birdcall clips.</li>\n</ul>\n<pre><code>call_paths_list = call_df[[\"paths\", \"pred_code\"]].values\nnocall_paths_list = nocall_df.paths.values\n\ndef create_stage3_cv(index):\n    k = random.choice([1,2,3,4,5])\n    nocalls = random.choises(nocall_paths_list, k=k)\n    calls = random.choises(call_paths_list, k=5-k)\n    audio_list = []\n    code_list = []\n    for f in nocalls:\n        y, _ = sf.read(f)\n        audio_list.append(y)\n    for l in calls:\n        path = l[0]\n        code = l[1]\n        y, _ = sf.read(path)\n        audio_list.append(y)\n        code_list.append(code)\n    random.shuffle(audio_list)\n    audio_cat = np.concatenate(audio_list)\n    codes = \"_\".join(code_list)\n    sf.write(f\"{index}_{codes}.wav\", audio_cat, sample_rate=16000)\n\n_ = Parallel(n_jobs=8, backend=\"multiprocessing\")(\n    delayed(create_stage3_cv)(i) for i in tqdm(range(160000//5)))\n)\n</code></pre>\n<p>Ex: <code>10000_sagthr_normoc_gryfly.wav</code> in this file you find 3bird calls and 2noise as 5sec clip.<br>\nyou find the cv dataset at <a href=\"https://www.kaggle.com/gopidurgaprasad/birdsong-stage3-cv\" target=\"_blank\">link</a></p>\n<h3>[Stage3]</h3>\n<ul>\n<li>on top of <strong>stage1</strong> and <strong>stage2</strong> models we trained <strong>stage3</strong> model using <strong>1sec birdcalls</strong> and <strong>1sec noise</strong>.</li>\n<li>the training idea is very simple as same as <strong>cv</strong>.</li>\n<li>at dataloder time we are taking 20% of <strong>1sec noise</strong> clips and 80% of <strong>1sec birdcalls</strong> clips</li>\n</ul>\n<pre><code>if np.random.random() &gt; 0.2:\n    y, sr = sf.read(wav_path)\n    labels[BIRD_CODE[ebird_code]] = 1\nelse:\n    y, sr = sf.read(random.choice(self.noise_files))\n    labels[BIRD_CODE[ebird_code]] = 0\n</code></pre>\n<ul>\n<li>at each batch time, we did something like shuffle and stack, inspired from cut mix and mixup</li>\n<li>In each batch, we have 20% noise and 80% birdcalls shuffle them and concatenate.</li>\n</ul>\n<pre><code>def stack_up(x, y, use_cuda=True):\n    batch_size = x.size()[0]\n    if use_cuda:\n        index0 = torch.randperm(batch_size).cuda()\n        index1 = torch.randperm(batch_size).cuda()\n        index2 = torch.randperm(batch_size).cuda()\n        index3 = torch.randperm(batch_size).cuda()\n        index4 = torch.randperm(batch_size).cuda()\n    ind = random.choice([0,1,2,3,4])\n    if ind == 0:\n        mixed_x = x\n        mixed_y = y\n    elif ind == 1:\n        mixed_x = torch.cat([x, x[index1,  :]], dim=1)\n        mixed_y = y + y[index1,  :]\n    elif ind == 2:\n        mixed_x = torch.cat([x, x[index1,  :], x[index2]], dim=1)\n        mixed_y = y + y[index1,  :] + y[index2,  :]\n    elif ind == 3:\n        mixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :]], dim=1)\n        mixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :]\n    elif ind == 4:\n        mixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :], x[index4,  :]], dim=1)\n        mixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :] + y[index4,  :]\n    mixed_y = torch.clamp(mixed_y, min=0, m[](url)ax=1)\n    return mixed_x, mixed_y\n</code></pre>\n<ul>\n<li>for <strong>stage3</strong> model we mouniter row_f1 score from this <a href=\"https://www.kaggle.com/shonenkov/competition-metrics\" target=\"_blank\">notebook</a> </li>\n<li>for every epoch our <strong>cv</strong> increased, then we conclude that we are going to trust this <strong>cv</strong></li>\n<li>at the end the best <strong>cv</strong> <strong>row_f1</strong> score in between <strong>[0.90 - 0.95]</strong></li>\n<li>this all processes are done in the last 3days so we are managed to train up to 5 models.</li>\n<li>at the end we don't have time as well as submissions, so we did a simple average on 5 models and using a simple threshold 0.5</li>\n<li>our average <strong>row_f1</strong> score is 0.94+ on 5 models.</li>\n</ul>\n<h3>[BirdSong North America Set]</h3>\n<ul>\n<li>for some folds in  <strong>stage3</strong> we are only trained on North America birds it improves our <strong>cv</strong></li>\n<li>you can find North America bird files in this <a href=\"https://www.kaggle.com/seshurajup/birdsong-north-america-set-stage-3\" target=\"_blank\">notebook</a> </li>\n</ul>\n<h3>[Stage3 Augmentations]</h3>\n<pre><code>import audiomentations as A\n\naugmenter = A.Compose([\n    A.AddGaussianNoise(p=0.3),\n    A.AddGaussianSNR(p=0.3),\n    A.AddBackgroundNoise(\"stage1_1sec_sudo_noise/\", p-0.5),\n    A.Normalize(p=0.2),\n    A.Gain(p=0.2)\n])\n</code></pre>\n<h3>[Stage3 Ensamble]</h3>\n<ul>\n<li>we trained our stage3 models 2dyas before the competition ending so we are managed to train 5 different models.</li>\n<li>1. <code>Cnn14_16k</code></li>\n<li>2. <code>resnest50d</code></li>\n<li>3. <code>efficientnet-03</code></li>\n<li>4. <code>efficientnet-04</code></li>\n<li>5. <code>efficientnet-05</code></li>\n<li>we did a simple average of those 5-models with simple threshold <code>0.5</code> our <em>cv</em> <code>0.94+</code> on LB: <code>0.533</code></li>\n<li>in the end, we satisfied our selves and trust our <strong>CV</strong> and we know we need to predict a mixed bird calls.</li>\n<li>so we selected as our final model and it gives Private LB: 0.632 </li>\n<li>we frankly saying that we are not able to beat the public LB score but we trusted our cv and training process</li>\n<li>that brings us 21st place in Private LB.</li>\n</ul>\n<blockquote>\n  <p>inference notebook : <a href=\"https://www.kaggle.com/gopidurgaprasad/birdcall-stage3-final\" target=\"_blank\">link</a></p>\n</blockquote>",
  "messages": [
    {
      "id": "1012192",
      "postDate": "09/16/2020 00:48:53",
      "content": "<p>First of all, I would like to thank kaggle and the organizers for hosting such an interesting competition.</p>\n<p>thanks for my wonderful team <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> and <a href=\"https://www.kaggle.com/tarique7\" target=\"_blank\">@tarique7</a>. this was great learning and collaboration.</p>\n<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> and I worked hard for this from the last 20days</p>\n<h4>[Big Shake-up]</h4>\n<ul>\n<li>LB : 0.533 ---&gt; Private LB : 0.632</li>\n<li>From 364th place ---&gt; 21st place</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2058044%2F0725e3d4482b444098f50d6df08441ef%2FScreenshot%20from%202020-09-16%2006-20-45.png?generation=1600217490267914&amp;alt=media\" alt=\"\"></p>\n<h3>[Summary]</h3>\n<ul>\n<li>converted mp3 audio file into wav with sample_rate 16k, because of colab drive space and birds sing up to 12kHz</li>\n<li>we trained <strong>stage1</strong> model based on 5sec random clip.</li>\n<li>It gives average of 0.65+ AMP on 5-fold</li>\n<li>we created 1sec audio clips dataset.</li>\n<li>predicted 5fold <strong>stage1</strong> models on 1sec dataset.</li>\n<li>selected <code>df[(df.ebird_code == df.pred_code) &amp; (df.pred_prob &gt;= 0.5)]</code> selected those 1sec clips for <strong>stage3</strong></li>\n<li>we trained <strong>stage2</strong> model on topof <strong>stage1</strong> using public data</li>\n<li>again predicted <strong>stage1</strong> and <strong>stage2</strong> models on public 1sec clips</li>\n<li>again select <code>df2[(df2.ebird_code == df2.pred_code) &amp; (df2.pred_prob &gt;= 0.5)]</code>selected those 1sec clips for <strong>stage3</strong></li>\n<li><strong>stage3</strong> dataset is <code>stage3_df = df.append(df2)</code> we endup with 612K 1sec clips with sudo labels</li>\n<li>we created 1sec noise clips using PANN <strong>Cnn14_16k</strong> model.</li>\n<li>we predict <strong>Cnn14_16k</strong> model on 1sec dataset and select some noise labels from PANN labels </li>\n<li>those labels are <code>['Silence', 'White noise', 'Vachical', 'Speech', 'Pink noise', 'Tick-tock', 'Wind noise (microphone)','Stream','Raindrop','Wind','Rain', ...... ]</code> selected those labels as noise labels</li>\n<li>based on those noise labels and <strong>stage1</strong> model predicted probabilities we select 1sec noise clips data.</li>\n<li><code>noise_df = df[(df.filename.isin(noise_labels_df.filename)) &amp; df.pred_prob &lt; 0.4]</code></li>\n<li>we end up with 103k noise 1sec clips you find dataset <a href=\"https://www.kaggle.com/gopidurgaprasad/birdsong-stage1-1sec-sudo-noise\" target=\"_blank\">link</a></li>\n<li>we know in this competition our main goal is to predict a mix of bird calls.</li>\n<li>now the main part, at the end we have <strong>stage1</strong>, <strong>stage2</strong> models, <strong>1sec</strong> dataset with Sudo labels, and <strong>1sec noise data</strong>. we trained the <strong>stage3</strong> model using all of those.</li>\n<li>now in front of us, we need to build <strong>CV</strong> and train a model that more reliable on predicting a mix of bird calls.</li>\n</ul>\n<h3>[CV]</h3>\n<ul>\n<li>we created a <strong>cv</strong> based on <strong>1sec bird calls</strong> and <strong>1sec noise data</strong></li>\n<li>In the end, we need to predict for <strong>5sec</strong> so we take 5 random birdcalls and noise stack them and give labels based on birdcall clips.</li>\n</ul>\n<pre><code>call_paths_list = call_df[[\"paths\", \"pred_code\"]].values\nnocall_paths_list = nocall_df.paths.values\n\ndef create_stage3_cv(index):\n    k = random.choice([1,2,3,4,5])\n    nocalls = random.choises(nocall_paths_list, k=k)\n    calls = random.choises(call_paths_list, k=5-k)\n    audio_list = []\n    code_list = []\n    for f in nocalls:\n        y, _ = sf.read(f)\n        audio_list.append(y)\n    for l in calls:\n        path = l[0]\n        code = l[1]\n        y, _ = sf.read(path)\n        audio_list.append(y)\n        code_list.append(code)\n    random.shuffle(audio_list)\n    audio_cat = np.concatenate(audio_list)\n    codes = \"_\".join(code_list)\n    sf.write(f\"{index}_{codes}.wav\", audio_cat, sample_rate=16000)\n\n_ = Parallel(n_jobs=8, backend=\"multiprocessing\")(\n    delayed(create_stage3_cv)(i) for i in tqdm(range(160000//5)))\n)\n</code></pre>\n<p>Ex: <code>10000_sagthr_normoc_gryfly.wav</code> in this file you find 3bird calls and 2noise as 5sec clip.<br>\nyou find the cv dataset at <a href=\"https://www.kaggle.com/gopidurgaprasad/birdsong-stage3-cv\" target=\"_blank\">link</a></p>\n<h3>[Stage3]</h3>\n<ul>\n<li>on top of <strong>stage1</strong> and <strong>stage2</strong> models we trained <strong>stage3</strong> model using <strong>1sec birdcalls</strong> and <strong>1sec noise</strong>.</li>\n<li>the training idea is very simple as same as <strong>cv</strong>.</li>\n<li>at dataloder time we are taking 20% of <strong>1sec noise</strong> clips and 80% of <strong>1sec birdcalls</strong> clips</li>\n</ul>\n<pre><code>if np.random.random() &gt; 0.2:\n    y, sr = sf.read(wav_path)\n    labels[BIRD_CODE[ebird_code]] = 1\nelse:\n    y, sr = sf.read(random.choice(self.noise_files))\n    labels[BIRD_CODE[ebird_code]] = 0\n</code></pre>\n<ul>\n<li>at each batch time, we did something like shuffle and stack, inspired from cut mix and mixup</li>\n<li>In each batch, we have 20% noise and 80% birdcalls shuffle them and concatenate.</li>\n</ul>\n<pre><code>def stack_up(x, y, use_cuda=True):\n    batch_size = x.size()[0]\n    if use_cuda:\n        index0 = torch.randperm(batch_size).cuda()\n        index1 = torch.randperm(batch_size).cuda()\n        index2 = torch.randperm(batch_size).cuda()\n        index3 = torch.randperm(batch_size).cuda()\n        index4 = torch.randperm(batch_size).cuda()\n    ind = random.choice([0,1,2,3,4])\n    if ind == 0:\n        mixed_x = x\n        mixed_y = y\n    elif ind == 1:\n        mixed_x = torch.cat([x, x[index1,  :]], dim=1)\n        mixed_y = y + y[index1,  :]\n    elif ind == 2:\n        mixed_x = torch.cat([x, x[index1,  :], x[index2]], dim=1)\n        mixed_y = y + y[index1,  :] + y[index2,  :]\n    elif ind == 3:\n        mixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :]], dim=1)\n        mixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :]\n    elif ind == 4:\n        mixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :], x[index4,  :]], dim=1)\n        mixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :] + y[index4,  :]\n    mixed_y = torch.clamp(mixed_y, min=0, m[](url)ax=1)\n    return mixed_x, mixed_y\n</code></pre>\n<ul>\n<li>for <strong>stage3</strong> model we mouniter row_f1 score from this <a href=\"https://www.kaggle.com/shonenkov/competition-metrics\" target=\"_blank\">notebook</a> </li>\n<li>for every epoch our <strong>cv</strong> increased, then we conclude that we are going to trust this <strong>cv</strong></li>\n<li>at the end the best <strong>cv</strong> <strong>row_f1</strong> score in between <strong>[0.90 - 0.95]</strong></li>\n<li>this all processes are done in the last 3days so we are managed to train up to 5 models.</li>\n<li>at the end we don't have time as well as submissions, so we did a simple average on 5 models and using a simple threshold 0.5</li>\n<li>our average <strong>row_f1</strong> score is 0.94+ on 5 models.</li>\n</ul>\n<h3>[BirdSong North America Set]</h3>\n<ul>\n<li>for some folds in  <strong>stage3</strong> we are only trained on North America birds it improves our <strong>cv</strong></li>\n<li>you can find North America bird files in this <a href=\"https://www.kaggle.com/seshurajup/birdsong-north-america-set-stage-3\" target=\"_blank\">notebook</a> </li>\n</ul>\n<h3>[Stage3 Augmentations]</h3>\n<pre><code>import audiomentations as A\n\naugmenter = A.Compose([\n    A.AddGaussianNoise(p=0.3),\n    A.AddGaussianSNR(p=0.3),\n    A.AddBackgroundNoise(\"stage1_1sec_sudo_noise/\", p-0.5),\n    A.Normalize(p=0.2),\n    A.Gain(p=0.2)\n])\n</code></pre>\n<h3>[Stage3 Ensamble]</h3>\n<ul>\n<li>we trained our stage3 models 2dyas before the competition ending so we are managed to train 5 different models.</li>\n<li>1. <code>Cnn14_16k</code></li>\n<li>2. <code>resnest50d</code></li>\n<li>3. <code>efficientnet-03</code></li>\n<li>4. <code>efficientnet-04</code></li>\n<li>5. <code>efficientnet-05</code></li>\n<li>we did a simple average of those 5-models with simple threshold <code>0.5</code> our <em>cv</em> <code>0.94+</code> on LB: <code>0.533</code></li>\n<li>in the end, we satisfied our selves and trust our <strong>CV</strong> and we know we need to predict a mixed bird calls.</li>\n<li>so we selected as our final model and it gives Private LB: 0.632 </li>\n<li>we frankly saying that we are not able to beat the public LB score but we trusted our cv and training process</li>\n<li>that brings us 21st place in Private LB.</li>\n</ul>\n<blockquote>\n  <p>inference notebook : <a href=\"https://www.kaggle.com/gopidurgaprasad/birdcall-stage3-final\" target=\"_blank\">link</a></p>\n</blockquote>",
      "rawMarkdown": "First of all, I would like to thank kaggle and the organizers for hosting such an interesting competition.\n\nthanks for my wonderful team @seshurajup and @tarique7. this was great learning and collaboration.\n\n@seshurajup and I worked hard for this from the last 20days\n\n#### [Big Shake-up]\n- LB : 0.533 ---> Private LB : 0.632\n- From 364th place ---> 21st place\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2058044%2F0725e3d4482b444098f50d6df08441ef%2FScreenshot%20from%202020-09-16%2006-20-45.png?generation=1600217490267914&alt=media)\n\n\n### [Summary]\n\n- converted mp3 audio file into wav with sample_rate 16k, because of colab drive space and birds sing up to 12kHz\n- we trained **stage1** model based on 5sec random clip.\n- It gives average of 0.65+ AMP on 5-fold\n- we created 1sec audio clips dataset.\n- predicted 5fold **stage1** models on 1sec dataset.\n- selected `df[(df.ebird_code == df.pred_code) & (df.pred_prob >= 0.5)]` selected those 1sec clips for **stage3**\n-  we trained **stage2** model on topof **stage1** using public data\n- again predicted **stage1** and **stage2** models on public 1sec clips\n- again select `df2[(df2.ebird_code == df2.pred_code) & (df2.pred_prob >= 0.5)]`selected those 1sec clips for **stage3**\n- **stage3** dataset is `stage3_df = df.append(df2)` we endup with 612K 1sec clips with sudo labels\n- we created 1sec noise clips using PANN **Cnn14_16k** model.\n- we predict **Cnn14_16k** model on 1sec dataset and select some noise labels from PANN labels \n- those labels are `['Silence', 'White noise', 'Vachical', 'Speech', 'Pink noise', 'Tick-tock', 'Wind noise (microphone)','Stream','Raindrop','Wind','Rain', ...... ]` selected those labels as noise labels\n- based on those noise labels and **stage1** model predicted probabilities we select 1sec noise clips data.\n- `noise_df = df[(df.filename.isin(noise_labels_df.filename)) & df.pred_prob < 0.4]`\n- we end up with 103k noise 1sec clips you find dataset [link](https://www.kaggle.com/gopidurgaprasad/birdsong-stage1-1sec-sudo-noise)\n- we know in this competition our main goal is to predict a mix of bird calls.\n- now the main part, at the end we have **stage1**, **stage2** models, **1sec** dataset with Sudo labels, and **1sec noise data**. we trained the **stage3** model using all of those.\n- now in front of us, we need to build **CV** and train a model that more reliable on predicting a mix of bird calls.\n\n### [CV]\n- we created a **cv** based on **1sec bird calls** and **1sec noise data**\n- In the end, we need to predict for **5sec** so we take 5 random birdcalls and noise stack them and give labels based on birdcall clips.\n\n```\ncall_paths_list = call_df[[\"paths\", \"pred_code\"]].values\nnocall_paths_list = nocall_df.paths.values\n\ndef create_stage3_cv(index):\n    k = random.choice([1,2,3,4,5])\n    nocalls = random.choises(nocall_paths_list, k=k)\n    calls = random.choises(call_paths_list, k=5-k)\n    audio_list = []\n    code_list = []\n    for f in nocalls:\n        y, _ = sf.read(f)\n        audio_list.append(y)\n    for l in calls:\n        path = l[0]\n        code = l[1]\n        y, _ = sf.read(path)\n        audio_list.append(y)\n        code_list.append(code)\n    random.shuffle(audio_list)\n    audio_cat = np.concatenate(audio_list)\n    codes = \"_\".join(code_list)\n    sf.write(f\"{index}_{codes}.wav\", audio_cat, sample_rate=16000)\n\n_ = Parallel(n_jobs=8, backend=\"multiprocessing\")(\n    delayed(create_stage3_cv)(i) for i in tqdm(range(160000//5)))\n)\n```\n\nEx: `10000_sagthr_normoc_gryfly.wav` in this file you find 3bird calls and 2noise as 5sec clip.\nyou find the cv dataset at [link](https://www.kaggle.com/gopidurgaprasad/birdsong-stage3-cv)\n\n### [Stage3]\n- on top of **stage1** and **stage2** models we trained **stage3** model using **1sec birdcalls** and **1sec noise**.\n- the training idea is very simple as same as **cv**.\n- at dataloder time we are taking 20% of **1sec noise** clips and 80% of **1sec birdcalls** clips\n```\nif np.random.random() > 0.2:\n\ty, sr = sf.read(wav_path)\n\tlabels[BIRD_CODE[ebird_code]] = 1\nelse:\n\ty, sr = sf.read(random.choice(self.noise_files))\n\tlabels[BIRD_CODE[ebird_code]] = 0\n```\n- at each batch time, we did something like shuffle and stack, inspired from cut mix and mixup\n- In each batch, we have 20% noise and 80% birdcalls shuffle them and concatenate.\n```\ndef stack_up(x, y, use_cuda=True):\n\tbatch_size = x.size()[0]\n\tif use_cuda:\n\t\tindex0 = torch.randperm(batch_size).cuda()\n\t\tindex1 = torch.randperm(batch_size).cuda()\n\t\tindex2 = torch.randperm(batch_size).cuda()\n\t\tindex3 = torch.randperm(batch_size).cuda()\n\t\tindex4 = torch.randperm(batch_size).cuda()\n\tind = random.choice([0,1,2,3,4])\n\tif ind == 0:\n\t\tmixed_x = x\n\t\tmixed_y = y\n\telif ind == 1:\n\t\tmixed_x = torch.cat([x, x[index1,  :]], dim=1)\n\t\tmixed_y = y + y[index1,  :]\n\telif ind == 2:\n\t\tmixed_x = torch.cat([x, x[index1,  :], x[index2]], dim=1)\n\t\tmixed_y = y + y[index1,  :] + y[index2,  :]\n\telif ind == 3:\n\t\tmixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :]], dim=1)\n\t\tmixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :]\n\telif ind == 4:\n\t\tmixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :], x[index4,  :]], dim=1)\n\t\tmixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :] + y[index4,  :]\n\tmixed_y = torch.clamp(mixed_y, min=0, m[](url)ax=1)\n\treturn mixed_x, mixed_y\n```\n- for **stage3** model we mouniter row_f1 score from this [notebook](https://www.kaggle.com/shonenkov/competition-metrics) \n- for every epoch our **cv** increased, then we conclude that we are going to trust this **cv**\n- at the end the best **cv** **row_f1** score in between **[0.90 - 0.95]**\n- this all processes are done in the last 3days so we are managed to train up to 5 models.\n- at the end we don't have time as well as submissions, so we did a simple average on 5 models and using a simple threshold 0.5\n- our average **row_f1** score is 0.94+ on 5 models.\n### [BirdSong North America Set]\n- for some folds in  **stage3** we are only trained on North America birds it improves our **cv**\n- you can find North America bird files in this [notebook](https://www.kaggle.com/seshurajup/birdsong-north-america-set-stage-3) \n\n### [Stage3 Augmentations]\n```\nimport audiomentations as A\n\naugmenter = A.Compose([\n\tA.AddGaussianNoise(p=0.3),\n\tA.AddGaussianSNR(p=0.3),\n\tA.AddBackgroundNoise(\"stage1_1sec_sudo_noise/\", p-0.5),\n\tA.Normalize(p=0.2),\n\tA.Gain(p=0.2)\n])\n```\n\n### [Stage3 Ensamble]\n- we trained our stage3 models 2dyas before the competition ending so we are managed to train 5 different models.\n- 1. `Cnn14_16k`\n- 2. `resnest50d`\n- 3. `efficientnet-03`\n- 4. `efficientnet-04`\n- 5. `efficientnet-05`\n- we did a simple average of those 5-models with simple threshold `0.5` our *cv* `0.94+` on LB: `0.533`\n- in the end, we satisfied our selves and trust our **CV** and we know we need to predict a mixed bird calls.\n- so we selected as our final model and it gives Private LB: 0.632 \n- we frankly saying that we are not able to beat the public LB score but we trusted our cv and training process\n- that brings us 21st place in Private LB.\n\n> inference notebook : [link](https://www.kaggle.com/gopidurgaprasad/birdcall-stage3-final)",
      "votes": null
    },
    {
      "id": "1012195",
      "postDate": "09/16/2020 00:50:19",
      "content": "<p>Thanks for participating in our competition! Looking forward to hearing more about your training process. </p>",
      "rawMarkdown": "Thanks for participating in our competition! Looking forward to hearing more about your training process.",
      "votes": null
    },
    {
      "id": "1012242",
      "postDate": "09/16/2020 01:36:46",
      "content": "<p>How can you sleep submitting something with public score 0.533 :)  Big congratulations !<br>\nThe birds sing up to 12kHz and you bandlimit to 8kHz.  Any reason you are confident this was ok ?  </p>",
      "rawMarkdown": "How can you sleep submitting something with public score 0.533 :)  Big congratulations !\nThe birds sing up to 12kHz and you bandlimit to 8kHz.  Any reason you are confident this was ok ?",
      "votes": null
    },
    {
      "id": "1012246",
      "postDate": "09/16/2020 01:48:07",
      "content": "<p><a href=\"https://www.kaggle.com/nyleve\" target=\"_blank\">@nyleve</a> Thank you so much 😁</p>\n<p>Very simple we trust our CV as well as training process. We alredy know we need to predict mix of bird calls.<br>\n16k because of disk space in colab, we are using colab for training. </p>\n<p>Don't bothered about 27% data…… Care about 73% Private data </p>",
      "rawMarkdown": "nyleve Thank you so much 😁\n\nVery simple we trust our CV as well as training process. We alredy know we need to predict mix of bird calls.\n16k because of disk space in colab, we are using colab for training. \n\nDon't bothered about 27% data...... Care about 73% Private data",
      "votes": null
    },
    {
      "id": "1012483",
      "postDate": "09/16/2020 05:47:48",
      "content": "<p>Congrats.</p>\n<p>My think is right (<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183015#1011029\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183015#1011029</a>)</p>",
      "rawMarkdown": "Congrats.\n\nMy think is right (https://www.kaggle.com/c/birdsong-recognition/discussion/183015#1011029)",
      "votes": null
    },
    {
      "id": "1012503",
      "postDate": "09/16/2020 05:57:02",
      "content": "<p><a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> thank you 😍</p>\n<p>absolutely </p>",
      "rawMarkdown": "truonghoang thank you 😍\n\nabsolutely",
      "votes": null
    },
    {
      "id": "1019236",
      "postDate": "09/20/2020 09:55:17",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "1019327",
      "postDate": "09/20/2020 11:02:48",
      "content": "<p><a href=\"https://www.kaggle.com/deepchatterjeevns\" target=\"_blank\">@deepchatterjeevns</a></p>\n<p>Thank you 😊</p>",
      "rawMarkdown": "deepchatterjeevns\n\nThank you 😊",
      "votes": null
    },
    {
      "id": "1020685",
      "postDate": "09/21/2020 10:57:43",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    },
    {
      "id": "1020741",
      "postDate": "09/21/2020 11:40:02",
      "content": "<p><a href=\"https://www.kaggle.com/tyadav\" target=\"_blank\">@tyadav</a> thank you 😊 </p>",
      "rawMarkdown": "tyadav thank you 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012195,
      "author_name": "holgerklinck",
      "author_url": "",
      "post_date": "09/16/2020 00:50:19",
      "content": "<p>Thanks for participating in our competition! Looking forward to hearing more about your training process. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012242,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "09/16/2020 01:36:46",
      "content": "<p>How can you sleep submitting something with public score 0.533 :)  Big congratulations !<br>\nThe birds sing up to 12kHz and you bandlimit to 8kHz.  Any reason you are confident this was ok ?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1012246,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/16/2020 01:48:07",
          "content": "<p><a href=\"https://www.kaggle.com/nyleve\" target=\"_blank\">@nyleve</a> Thank you so much 😁</p>\n<p>Very simple we trust our CV as well as training process. We alredy know we need to predict mix of bird calls.<br>\n16k because of disk space in colab, we are using colab for training. </p>\n<p>Don't bothered about 27% data…… Care about 73% Private data </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012483,
      "author_name": "truonghoang",
      "author_url": "",
      "post_date": "09/16/2020 05:47:48",
      "content": "<p>Congrats.</p>\n<p>My think is right (<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183015#1011029\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183015#1011029</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012503,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/16/2020 05:57:02",
          "content": "<p><a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> thank you 😍</p>\n<p>absolutely </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1019236,
      "author_name": "deepchatterjeevns",
      "author_url": "",
      "post_date": "09/20/2020 09:55:17",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1019327,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/20/2020 11:02:48",
          "content": "<p><a href=\"https://www.kaggle.com/deepchatterjeevns\" target=\"_blank\">@deepchatterjeevns</a></p>\n<p>Thank you 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1020685,
      "author_name": "tyadav",
      "author_url": "",
      "post_date": "09/21/2020 10:57:43",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1020741,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/21/2020 11:40:02",
          "content": "<p><a href=\"https://www.kaggle.com/tyadav\" target=\"_blank\">@tyadav</a> thank you 😊 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1012192": "First of all, I would like to thank kaggle and the organizers for hosting such an interesting competition.\n\nthanks for my wonderful team @seshurajup and @tarique7. this was great learning and collaboration.\n\n@seshurajup and I worked hard for this from the last 20days\n\n#### [Big Shake-up]\n- LB : 0.533 ---> Private LB : 0.632\n- From 364th place ---> 21st place\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2058044%2F0725e3d4482b444098f50d6df08441ef%2FScreenshot%20from%202020-09-16%2006-20-45.png?generation=1600217490267914&alt=media)\n\n\n### [Summary]\n\n- converted mp3 audio file into wav with sample_rate 16k, because of colab drive space and birds sing up to 12kHz\n- we trained **stage1** model based on 5sec random clip.\n- It gives average of 0.65+ AMP on 5-fold\n- we created 1sec audio clips dataset.\n- predicted 5fold **stage1** models on 1sec dataset.\n- selected `df[(df.ebird_code == df.pred_code) & (df.pred_prob >= 0.5)]` selected those 1sec clips for **stage3**\n-  we trained **stage2** model on topof **stage1** using public data\n- again predicted **stage1** and **stage2** models on public 1sec clips\n- again select `df2[(df2.ebird_code == df2.pred_code) & (df2.pred_prob >= 0.5)]`selected those 1sec clips for **stage3**\n- **stage3** dataset is `stage3_df = df.append(df2)` we endup with 612K 1sec clips with sudo labels\n- we created 1sec noise clips using PANN **Cnn14_16k** model.\n- we predict **Cnn14_16k** model on 1sec dataset and select some noise labels from PANN labels \n- those labels are `['Silence', 'White noise', 'Vachical', 'Speech', 'Pink noise', 'Tick-tock', 'Wind noise (microphone)','Stream','Raindrop','Wind','Rain', ...... ]` selected those labels as noise labels\n- based on those noise labels and **stage1** model predicted probabilities we select 1sec noise clips data.\n- `noise_df = df[(df.filename.isin(noise_labels_df.filename)) & df.pred_prob < 0.4]`\n- we end up with 103k noise 1sec clips you find dataset [link](https://www.kaggle.com/gopidurgaprasad/birdsong-stage1-1sec-sudo-noise)\n- we know in this competition our main goal is to predict a mix of bird calls.\n- now the main part, at the end we have **stage1**, **stage2** models, **1sec** dataset with Sudo labels, and **1sec noise data**. we trained the **stage3** model using all of those.\n- now in front of us, we need to build **CV** and train a model that more reliable on predicting a mix of bird calls.\n\n### [CV]\n- we created a **cv** based on **1sec bird calls** and **1sec noise data**\n- In the end, we need to predict for **5sec** so we take 5 random birdcalls and noise stack them and give labels based on birdcall clips.\n\n```\ncall_paths_list = call_df[[\"paths\", \"pred_code\"]].values\nnocall_paths_list = nocall_df.paths.values\n\ndef create_stage3_cv(index):\n    k = random.choice([1,2,3,4,5])\n    nocalls = random.choises(nocall_paths_list, k=k)\n    calls = random.choises(call_paths_list, k=5-k)\n    audio_list = []\n    code_list = []\n    for f in nocalls:\n        y, _ = sf.read(f)\n        audio_list.append(y)\n    for l in calls:\n        path = l[0]\n        code = l[1]\n        y, _ = sf.read(path)\n        audio_list.append(y)\n        code_list.append(code)\n    random.shuffle(audio_list)\n    audio_cat = np.concatenate(audio_list)\n    codes = \"_\".join(code_list)\n    sf.write(f\"{index}_{codes}.wav\", audio_cat, sample_rate=16000)\n\n_ = Parallel(n_jobs=8, backend=\"multiprocessing\")(\n    delayed(create_stage3_cv)(i) for i in tqdm(range(160000//5)))\n)\n```\n\nEx: `10000_sagthr_normoc_gryfly.wav` in this file you find 3bird calls and 2noise as 5sec clip.\nyou find the cv dataset at [link](https://www.kaggle.com/gopidurgaprasad/birdsong-stage3-cv)\n\n### [Stage3]\n- on top of **stage1** and **stage2** models we trained **stage3** model using **1sec birdcalls** and **1sec noise**.\n- the training idea is very simple as same as **cv**.\n- at dataloder time we are taking 20% of **1sec noise** clips and 80% of **1sec birdcalls** clips\n```\nif np.random.random() > 0.2:\n\ty, sr = sf.read(wav_path)\n\tlabels[BIRD_CODE[ebird_code]] = 1\nelse:\n\ty, sr = sf.read(random.choice(self.noise_files))\n\tlabels[BIRD_CODE[ebird_code]] = 0\n```\n- at each batch time, we did something like shuffle and stack, inspired from cut mix and mixup\n- In each batch, we have 20% noise and 80% birdcalls shuffle them and concatenate.\n```\ndef stack_up(x, y, use_cuda=True):\n\tbatch_size = x.size()[0]\n\tif use_cuda:\n\t\tindex0 = torch.randperm(batch_size).cuda()\n\t\tindex1 = torch.randperm(batch_size).cuda()\n\t\tindex2 = torch.randperm(batch_size).cuda()\n\t\tindex3 = torch.randperm(batch_size).cuda()\n\t\tindex4 = torch.randperm(batch_size).cuda()\n\tind = random.choice([0,1,2,3,4])\n\tif ind == 0:\n\t\tmixed_x = x\n\t\tmixed_y = y\n\telif ind == 1:\n\t\tmixed_x = torch.cat([x, x[index1,  :]], dim=1)\n\t\tmixed_y = y + y[index1,  :]\n\telif ind == 2:\n\t\tmixed_x = torch.cat([x, x[index1,  :], x[index2]], dim=1)\n\t\tmixed_y = y + y[index1,  :] + y[index2,  :]\n\telif ind == 3:\n\t\tmixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :]], dim=1)\n\t\tmixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :]\n\telif ind == 4:\n\t\tmixed_x = torch.cat([x, x[index1,  :], x[index2], x[index3,  :], x[index4,  :]], dim=1)\n\t\tmixed_y = y + y[index1,  :] + y[index2,  :] + y[index3,  :] + y[index4,  :]\n\tmixed_y = torch.clamp(mixed_y, min=0, m[](url)ax=1)\n\treturn mixed_x, mixed_y\n```\n- for **stage3** model we mouniter row_f1 score from this [notebook](https://www.kaggle.com/shonenkov/competition-metrics) \n- for every epoch our **cv** increased, then we conclude that we are going to trust this **cv**\n- at the end the best **cv** **row_f1** score in between **[0.90 - 0.95]**\n- this all processes are done in the last 3days so we are managed to train up to 5 models.\n- at the end we don't have time as well as submissions, so we did a simple average on 5 models and using a simple threshold 0.5\n- our average **row_f1** score is 0.94+ on 5 models.\n### [BirdSong North America Set]\n- for some folds in  **stage3** we are only trained on North America birds it improves our **cv**\n- you can find North America bird files in this [notebook](https://www.kaggle.com/seshurajup/birdsong-north-america-set-stage-3) \n\n### [Stage3 Augmentations]\n```\nimport audiomentations as A\n\naugmenter = A.Compose([\n\tA.AddGaussianNoise(p=0.3),\n\tA.AddGaussianSNR(p=0.3),\n\tA.AddBackgroundNoise(\"stage1_1sec_sudo_noise/\", p-0.5),\n\tA.Normalize(p=0.2),\n\tA.Gain(p=0.2)\n])\n```\n\n### [Stage3 Ensamble]\n- we trained our stage3 models 2dyas before the competition ending so we are managed to train 5 different models.\n- 1. `Cnn14_16k`\n- 2. `resnest50d`\n- 3. `efficientnet-03`\n- 4. `efficientnet-04`\n- 5. `efficientnet-05`\n- we did a simple average of those 5-models with simple threshold `0.5` our *cv* `0.94+` on LB: `0.533`\n- in the end, we satisfied our selves and trust our **CV** and we know we need to predict a mixed bird calls.\n- so we selected as our final model and it gives Private LB: 0.632 \n- we frankly saying that we are not able to beat the public LB score but we trusted our cv and training process\n- that brings us 21st place in Private LB.\n\n> inference notebook : [link](https://www.kaggle.com/gopidurgaprasad/birdcall-stage3-final)",
    "1012195": "Thanks for participating in our competition! Looking forward to hearing more about your training process.",
    "1012242": "How can you sleep submitting something with public score 0.533 :)  Big congratulations !\nThe birds sing up to 12kHz and you bandlimit to 8kHz.  Any reason you are confident this was ok ?",
    "1012246": "nyleve Thank you so much 😁\n\nVery simple we trust our CV as well as training process. We alredy know we need to predict mix of bird calls.\n16k because of disk space in colab, we are using colab for training. \n\nDon't bothered about 27% data...... Care about 73% Private data",
    "1012483": "Congrats.\n\nMy think is right (https://www.kaggle.com/c/birdsong-recognition/discussion/183015#1011029)",
    "1012503": "truonghoang thank you 😍\n\nabsolutely",
    "1019236": "Congratulations!",
    "1019327": "deepchatterjeevns\n\nThank you 😊",
    "1020685": "Congratulations!!",
    "1020741": "tyadav thank you 😊"
  },
  "source": "meta"
}