{
  "id": 327187,
  "title": "6th place solution (human-in-the-loop)",
  "url": "/competitions/birdclef-2022/writeups/shinmura0-6th-place-solution-human-in-the-loop",
  "author_name": "",
  "post_date": "2022-05-26T03:04:45.913145600Z",
  "votes": 38,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First I want to thank this competition hosts and the Kaggle team for organizing such a interesting competition. And thank you to all the Kagglers.</p>\n<h2>Overview</h2>\n<ul>\n<li>Only using SED model</li>\n<li>Make clean data with <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319813\" target=\"_blank\">human-in-the-loop</a> for weak label</li>\n<li>Heavy post-processing (public LB 0.79 → 0.84)</li>\n</ul>\n<p>Perhaps my SED is the same model as yours.<br>\n<strong>Clean data</strong> and <strong>post-processing</strong> is important in my solution.</p>\n<h2>Clean data (important)</h2>\n<p>I made <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/239911\" target=\"_blank\">annotations</a> in BirdClef2021. And this is effective. But hand labeling is time consuming.<br>\nInstead of hand labeling, I used human-in-the-loop method in this competition.</p>\n<ul>\n<li>First, train SED with primary label</li>\n<li>Second, extract high confidence SED predictions in training data</li>\n<li>I listen these prediction and judge correct or not correct</li>\n<li>Human verification (I answer yes or no)</li>\n</ul>\n<p>I made 2,000 clean data by this human-in-the-loop. These data contain only scored bird.</p>\n<h2>Training data</h2>\n<p>I used 3 type data.</p>\n<ul>\n<li>HIL_data (=human in the loop data)<ul>\n<li>contain <a href=\"https://www.kaggle.com/code/amandanavine/hawaiian-bird-species\" target=\"_blank\">external data</a></li></ul></li>\n<li>other_data (not scored bird data(about 130species))<ul>\n<li>initial 3sec audio</li>\n<li>The label is primary_label</li></ul></li>\n<li>psuedo_data (scored bird, and not contain HIL audio)<ul>\n<li>initial 3sec audio</li>\n<li>psuedo_label = 0.25primay_label + 0.25x1st_generation_model + 0.5x2nd_generation_model </li></ul></li>\n</ul>\n<p>And training data is below ratio.<br>\nThis is a best ratio.</p>\n<p>HIL_data : other_data : psuedo_data = 1 : 4 : 1</p>\n<h2>SED</h2>\n<ul>\n<li>Backbone: eca_nfnet_l0, dm_nfnet_f0</li>\n<li>Only using \"clipwise_output\" (training &amp; inference)</li>\n<li><a href=\"https://www.kaggle.com/code/shinmurashinmura/birdclef2022-basic-augmentation/notebook\" target=\"_blank\">Basic augment</a><ul>\n<li>Time shift</li>\n<li>Add pink noise and brown noise</li>\n<li>Mix other audio dataset (ESC-50: frog, rain, airplane, crackling_fire)</li></ul></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/307880\" target=\"_blank\">SpecAugment++</a> (mixing Mel-spec is ESC-50)</li>\n<li>Label smoothing (alpha=0.1)</li>\n<li>Optimize with Adam<ul>\n<li>lr=0.0001</li>\n<li>CosineAnealing (T=10)</li></ul></li>\n<li>Epoch:30 (about 14hours in Colaboratory)</li>\n<li>Loss: BCEWithLogitsLoss</li>\n<li>Input (train &amp; inference): 5sec</li>\n<li>STFT resolution:250x254<ul>\n<li>window_size: 1024</li>\n<li>hop_size: 630</li>\n<li>mel_bins: 250</li>\n<li>fmin: 50</li>\n<li>fmax: 14000</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Ensemble is a little impact in this competition</li>\n<li>I compared voting vs average.<ul>\n<li>Voting is good score a little.</li></ul></li>\n<li>Finally, I used voting with 10 models.</li>\n</ul>\n<h2>Post-Processing (important)</h2>\n<p>My Post-Processing is similar with <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326979\" target=\"_blank\">12th place solution</a>.</p>\n<h4>prediction time shift (=predition_TA)</h4>\n<p>prediction_now = now + 0.5next_5sec + 0.25next_10sec + 0.5previous_5sec + 0.25previous_10sec</p>\n<h4>TTA</h4>\n<p>Let t be the target time. I used 3 type inputs.</p>\n<ul>\n<li>[t,t+5]</li>\n<li>[t-1,t+4]</li>\n<li>[t+1,t+6]</li>\n</ul>\n<p>And these ouput is used in voting.</p>\n<h4>Threshold optimization (=ThreshO)</h4>\n<ul>\n<li>First, I tuned threhold like <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">this</a>. These threshold is constant.</li>\n<li>Second, I tuned threhold each species.</li>\n</ul>\n<h4>Threshold down (=ThershD)</h4>\n<p>If models detect a certain bird once in the 1 minute audio, I lower the threshold in the audio and infer it once more. </p>\n<pre><code>score = prediction(thresh) # 1min prediction\ndown_species = [\"hawhaw\", \"hawpet1\", \"maupar\", \"ercfra\", \"crehon\", \"puaioh\", \"hawcre\"]\n\nfor bird in down_species:\n    if np.max(score[bird]) &gt; thresh[bird]:\n        thresh[bird] = thresh[bird] - decrease\n\nscore = prediction(thresh) # 1min prediction once more\n</code></pre>\n<h2>Ablation study</h2>\n<table>\n<thead>\n<tr>\n<th>post-processing</th>\n<th><strong>Public</strong>/Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>not using</td>\n<td><strong>0.79</strong>/0.76</td>\n</tr>\n<tr>\n<td>predition_TA</td>\n<td><strong>0.78</strong>/0.75</td>\n</tr>\n<tr>\n<td>predition_TA + TTA</td>\n<td><strong>0.78</strong>/0.74</td>\n</tr>\n<tr>\n<td>predition_TA + TTA + ThreshO</td>\n<td><strong>0.82</strong>/0.77</td>\n</tr>\n<tr>\n<td>predition_TA + TTA + ThreshO + ThreshD</td>\n<td><strong>0.84</strong>/0.80</td>\n</tr>\n</tbody>\n</table>\n<h2>Not working</h2>\n<ul>\n<li>AST (equally SED)</li>\n<li>PCEN (equally Mel-spec)</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/307880\" target=\"_blank\">ImportantAug</a></li>\n<li>ArcFace as few-shot learning method</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243360\" target=\"_blank\">STFT Transformer</a></li>\n<li>Cooccurrence species with <a href=\"https://www.iucnredlist.org/species/22708583/128101101\" target=\"_blank\">these map</a></li>\n<li>Using BirdClef2021 dataset</li>\n</ul>",
  "messages": [
    {
      "id": "1801673",
      "postDate": "05/26/2022 03:04:45",
      "content": "<p>First I want to thank this competition hosts and the Kaggle team for organizing such a interesting competition. And thank you to all the Kagglers.</p>\n<h2>Overview</h2>\n<ul>\n<li>Only using SED model</li>\n<li>Make clean data with <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319813\" target=\"_blank\">human-in-the-loop</a> for weak label</li>\n<li>Heavy post-processing (public LB 0.79 → 0.84)</li>\n</ul>\n<p>Perhaps my SED is the same model as yours.<br>\n<strong>Clean data</strong> and <strong>post-processing</strong> is important in my solution.</p>\n<h2>Clean data (important)</h2>\n<p>I made <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/239911\" target=\"_blank\">annotations</a> in BirdClef2021. And this is effective. But hand labeling is time consuming.<br>\nInstead of hand labeling, I used human-in-the-loop method in this competition.</p>\n<ul>\n<li>First, train SED with primary label</li>\n<li>Second, extract high confidence SED predictions in training data</li>\n<li>I listen these prediction and judge correct or not correct</li>\n<li>Human verification (I answer yes or no)</li>\n</ul>\n<p>I made 2,000 clean data by this human-in-the-loop. These data contain only scored bird.</p>\n<h2>Training data</h2>\n<p>I used 3 type data.</p>\n<ul>\n<li>HIL_data (=human in the loop data)<ul>\n<li>contain <a href=\"https://www.kaggle.com/code/amandanavine/hawaiian-bird-species\" target=\"_blank\">external data</a></li></ul></li>\n<li>other_data (not scored bird data(about 130species))<ul>\n<li>initial 3sec audio</li>\n<li>The label is primary_label</li></ul></li>\n<li>psuedo_data (scored bird, and not contain HIL audio)<ul>\n<li>initial 3sec audio</li>\n<li>psuedo_label = 0.25primay_label + 0.25x1st_generation_model + 0.5x2nd_generation_model </li></ul></li>\n</ul>\n<p>And training data is below ratio.<br>\nThis is a best ratio.</p>\n<p>HIL_data : other_data : psuedo_data = 1 : 4 : 1</p>\n<h2>SED</h2>\n<ul>\n<li>Backbone: eca_nfnet_l0, dm_nfnet_f0</li>\n<li>Only using \"clipwise_output\" (training &amp; inference)</li>\n<li><a href=\"https://www.kaggle.com/code/shinmurashinmura/birdclef2022-basic-augmentation/notebook\" target=\"_blank\">Basic augment</a><ul>\n<li>Time shift</li>\n<li>Add pink noise and brown noise</li>\n<li>Mix other audio dataset (ESC-50: frog, rain, airplane, crackling_fire)</li></ul></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/307880\" target=\"_blank\">SpecAugment++</a> (mixing Mel-spec is ESC-50)</li>\n<li>Label smoothing (alpha=0.1)</li>\n<li>Optimize with Adam<ul>\n<li>lr=0.0001</li>\n<li>CosineAnealing (T=10)</li></ul></li>\n<li>Epoch:30 (about 14hours in Colaboratory)</li>\n<li>Loss: BCEWithLogitsLoss</li>\n<li>Input (train &amp; inference): 5sec</li>\n<li>STFT resolution:250x254<ul>\n<li>window_size: 1024</li>\n<li>hop_size: 630</li>\n<li>mel_bins: 250</li>\n<li>fmin: 50</li>\n<li>fmax: 14000</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Ensemble is a little impact in this competition</li>\n<li>I compared voting vs average.<ul>\n<li>Voting is good score a little.</li></ul></li>\n<li>Finally, I used voting with 10 models.</li>\n</ul>\n<h2>Post-Processing (important)</h2>\n<p>My Post-Processing is similar with <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326979\" target=\"_blank\">12th place solution</a>.</p>\n<h4>prediction time shift (=predition_TA)</h4>\n<p>prediction_now = now + 0.5next_5sec + 0.25next_10sec + 0.5previous_5sec + 0.25previous_10sec</p>\n<h4>TTA</h4>\n<p>Let t be the target time. I used 3 type inputs.</p>\n<ul>\n<li>[t,t+5]</li>\n<li>[t-1,t+4]</li>\n<li>[t+1,t+6]</li>\n</ul>\n<p>And these ouput is used in voting.</p>\n<h4>Threshold optimization (=ThreshO)</h4>\n<ul>\n<li>First, I tuned threhold like <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">this</a>. These threshold is constant.</li>\n<li>Second, I tuned threhold each species.</li>\n</ul>\n<h4>Threshold down (=ThershD)</h4>\n<p>If models detect a certain bird once in the 1 minute audio, I lower the threshold in the audio and infer it once more. </p>\n<pre><code>score = prediction(thresh) # 1min prediction\ndown_species = [\"hawhaw\", \"hawpet1\", \"maupar\", \"ercfra\", \"crehon\", \"puaioh\", \"hawcre\"]\n\nfor bird in down_species:\n    if np.max(score[bird]) &gt; thresh[bird]:\n        thresh[bird] = thresh[bird] - decrease\n\nscore = prediction(thresh) # 1min prediction once more\n</code></pre>\n<h2>Ablation study</h2>\n<table>\n<thead>\n<tr>\n<th>post-processing</th>\n<th><strong>Public</strong>/Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>not using</td>\n<td><strong>0.79</strong>/0.76</td>\n</tr>\n<tr>\n<td>predition_TA</td>\n<td><strong>0.78</strong>/0.75</td>\n</tr>\n<tr>\n<td>predition_TA + TTA</td>\n<td><strong>0.78</strong>/0.74</td>\n</tr>\n<tr>\n<td>predition_TA + TTA + ThreshO</td>\n<td><strong>0.82</strong>/0.77</td>\n</tr>\n<tr>\n<td>predition_TA + TTA + ThreshO + ThreshD</td>\n<td><strong>0.84</strong>/0.80</td>\n</tr>\n</tbody>\n</table>\n<h2>Not working</h2>\n<ul>\n<li>AST (equally SED)</li>\n<li>PCEN (equally Mel-spec)</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/307880\" target=\"_blank\">ImportantAug</a></li>\n<li>ArcFace as few-shot learning method</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243360\" target=\"_blank\">STFT Transformer</a></li>\n<li>Cooccurrence species with <a href=\"https://www.iucnredlist.org/species/22708583/128101101\" target=\"_blank\">these map</a></li>\n<li>Using BirdClef2021 dataset</li>\n</ul>",
      "rawMarkdown": "First I want to thank this competition hosts and the Kaggle team for organizing such a interesting competition. And thank you to all the Kagglers.\n\n## Overview\n- Only using SED model\n- Make clean data with [human-in-the-loop](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319813) for weak label\n- Heavy post-processing (public LB 0.79 &rarr; 0.84)\n\nPerhaps my SED is the same model as yours.\n**Clean data** and **post-processing** is important in my solution.\n\n## Clean data (important)\nI made [annotations](https://www.kaggle.com/competitions/birdclef-2021/discussion/239911) in BirdClef2021. And this is effective. But hand labeling is time consuming.\nInstead of hand labeling, I used human-in-the-loop method in this competition.\n- First, train SED with primary label\n- Second, extract high confidence SED predictions in training data\n- I listen these prediction and judge correct or not correct\n- Human verification (I answer yes or no)\n\nI made 2,000 clean data by this human-in-the-loop. These data contain only scored bird.\n\n## Training data\nI used 3 type data.\n- HIL_data (=human in the loop data)\n  - contain [external data](https://www.kaggle.com/code/amandanavine/hawaiian-bird-species)\n- other_data (not scored bird data(about 130species))\n  - initial 3sec audio\n  - The label is primary_label\n- psuedo_data (scored bird, and not contain HIL audio)\n  - initial 3sec audio\n  - psuedo_label = 0.25primay_label + 0.25x1st_generation_model + 0.5x2nd_generation_model \n\nAnd training data is below ratio.\nThis is a best ratio.\n\nHIL_data : other_data : psuedo_data = 1 : 4 : 1\n\n## SED\n- Backbone: eca_nfnet_l0, dm_nfnet_f0\n- Only using \"clipwise_output\" (training & inference)\n- [Basic augment](https://www.kaggle.com/code/shinmurashinmura/birdclef2022-basic-augmentation/notebook)\n  - Time shift\n  - Add pink noise and brown noise\n  - Mix other audio dataset (ESC-50: frog, rain, airplane, crackling_fire)\n- [SpecAugment++](https://www.kaggle.com/competitions/birdclef-2022/discussion/307880) (mixing Mel-spec is ESC-50)\n- Label smoothing (alpha=0.1)\n- Optimize with Adam\n  - lr=0.0001\n  - CosineAnealing (T=10)\n- Epoch:30 (about 14hours in Colaboratory)\n- Loss: BCEWithLogitsLoss\n- Input (train & inference): 5sec\n- STFT resolution:250x254\n  - window_size: 1024\n  - hop_size: 630\n  - mel_bins: 250\n  - fmin: 50\n  - fmax: 14000\n\n## Ensemble\n- Ensemble is a little impact in this competition\n- I compared voting vs average.\n  - Voting is good score a little.\n- Finally, I used voting with 10 models.\n\n## Post-Processing (important)\nMy Post-Processing is similar with [12th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/326979).\n \n#### prediction time shift (=predition_TA)\nprediction_now = now + 0.5next_5sec + 0.25next_10sec + 0.5previous_5sec + 0.25previous_10sec\n#### TTA\nLet t be the target time. I used 3 type inputs.\n- [t,t+5]\n- [t-1,t+4]\n- [t+1,t+6]\n\nAnd these ouput is used in voting.\n#### Threshold optimization (=ThreshO)\n- First, I tuned threhold like [this](https://www.kaggle.com/competitions/birdclef-2022/discussion/318999). These threshold is constant.\n- Second, I tuned threhold each species.\n#### Threshold down (=ThershD)\nIf models detect a certain bird once in the 1 minute audio, I lower the threshold in the audio and infer it once more. \n\n```\nscore = prediction(thresh) # 1min prediction\ndown_species = [\"hawhaw\", \"hawpet1\", \"maupar\", \"ercfra\", \"crehon\", \"puaioh\", \"hawcre\"]\n\nfor bird in down_species:\n    if np.max(score[bird]) > thresh[bird]:\n        thresh[bird] = thresh[bird] - decrease\n\nscore = prediction(thresh) # 1min prediction once more\n```\n\n## Ablation study\n|post-processing|**Public**/Private LB\n|---|---|\n|not using|**0.79**/0.76\n|predition_TA|**0.78**/0.75\n|predition_TA + TTA|**0.78**/0.74\n|predition_TA + TTA + ThreshO|**0.82**/0.77\n|predition_TA + TTA + ThreshO + ThreshD|**0.84**/0.80\n\n## Not working\n- AST (equally SED)\n- PCEN (equally Mel-spec)\n- [ImportantAug](https://www.kaggle.com/competitions/birdclef-2022/discussion/307880)\n- ArcFace as few-shot learning method\n- [STFT Transformer](https://www.kaggle.com/competitions/birdclef-2021/discussion/243360)\n- Cooccurrence species with [these map](https://www.iucnredlist.org/species/22708583/128101101)\n- Using BirdClef2021 dataset",
      "votes": null
    },
    {
      "id": "1801703",
      "postDate": "05/26/2022 04:13:19",
      "content": "<p>Congrats on your solo gold and competition master!<br>\nThank you for sharing your solution. How did you tune threshold for each species?</p>",
      "rawMarkdown": "Congrats on your solo gold and competition master!\nThank you for sharing your solution. How did you tune threshold for each species?",
      "votes": null
    },
    {
      "id": "1801870",
      "postDate": "05/26/2022 08:18:12",
      "content": "<p>It's trial and trial.<br>\nI ran one line at a time like below.</p>\n<pre><code>thresh[label_dic[\"skylar\"]] =  thresh[label_dic[\"skylar\"]] + 0.2\nthresh[label_dic[\"hawcre\"]] =  thresh[label_dic[\"hawcre\"]] + 0.2\nthresh[label_dic[\"aniani\"]] =  thresh[label_dic[\"aniani\"]] + 0.2\nthresh[label_dic[\"hawama\"]] =  thresh[label_dic[\"hawama\"]] + 0.2\n#thresh[label_dic[\"warwhe1\"]] =  thresh[label_dic[\"warwhe1\"]] + 0.2\nthresh[label_dic[\"yefcan\"]] =  thresh[label_dic[\"yefcan\"]] + 0.2\n#thresh[label_dic[\"jabwar\"]] =  thresh[label_dic[\"jabwar\"]] + 0.2\nthresh[label_dic[\"houfin\"]] =  thresh[label_dic[\"houfin\"]] + 0.2\nthresh[label_dic[\"apapan\"]] =  thresh[label_dic[\"apapan\"]] + 0.2\n</code></pre>\n<p>Mainly, the threshold for the majority species was raised.</p>",
      "rawMarkdown": "It's trial and trial.\nI ran one line at a time like below.\n\n```\nthresh[label_dic[\"skylar\"]] =  thresh[label_dic[\"skylar\"]] + 0.2\nthresh[label_dic[\"hawcre\"]] =  thresh[label_dic[\"hawcre\"]] + 0.2\nthresh[label_dic[\"aniani\"]] =  thresh[label_dic[\"aniani\"]] + 0.2\nthresh[label_dic[\"hawama\"]] =  thresh[label_dic[\"hawama\"]] + 0.2\n#thresh[label_dic[\"warwhe1\"]] =  thresh[label_dic[\"warwhe1\"]] + 0.2\nthresh[label_dic[\"yefcan\"]] =  thresh[label_dic[\"yefcan\"]] + 0.2\n#thresh[label_dic[\"jabwar\"]] =  thresh[label_dic[\"jabwar\"]] + 0.2\nthresh[label_dic[\"houfin\"]] =  thresh[label_dic[\"houfin\"]] + 0.2\nthresh[label_dic[\"apapan\"]] =  thresh[label_dic[\"apapan\"]] + 0.2\n```\n\nMainly, the threshold for the majority species was raised.",
      "votes": null
    },
    {
      "id": "1801903",
      "postDate": "05/26/2022 09:05:44",
      "content": "<p>Thanks. If the only way of tuning is to submit, it's hard.</p>",
      "rawMarkdown": "Thanks. If the only way of tuning is to submit, it's hard.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801703,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "05/26/2022 04:13:19",
      "content": "<p>Congrats on your solo gold and competition master!<br>\nThank you for sharing your solution. How did you tune threshold for each species?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801870,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "05/26/2022 08:18:12",
          "content": "<p>It's trial and trial.<br>\nI ran one line at a time like below.</p>\n<pre><code>thresh[label_dic[\"skylar\"]] =  thresh[label_dic[\"skylar\"]] + 0.2\nthresh[label_dic[\"hawcre\"]] =  thresh[label_dic[\"hawcre\"]] + 0.2\nthresh[label_dic[\"aniani\"]] =  thresh[label_dic[\"aniani\"]] + 0.2\nthresh[label_dic[\"hawama\"]] =  thresh[label_dic[\"hawama\"]] + 0.2\n#thresh[label_dic[\"warwhe1\"]] =  thresh[label_dic[\"warwhe1\"]] + 0.2\nthresh[label_dic[\"yefcan\"]] =  thresh[label_dic[\"yefcan\"]] + 0.2\n#thresh[label_dic[\"jabwar\"]] =  thresh[label_dic[\"jabwar\"]] + 0.2\nthresh[label_dic[\"houfin\"]] =  thresh[label_dic[\"houfin\"]] + 0.2\nthresh[label_dic[\"apapan\"]] =  thresh[label_dic[\"apapan\"]] + 0.2\n</code></pre>\n<p>Mainly, the threshold for the majority species was raised.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801903,
          "author_name": "shigemitsutomizawa",
          "author_url": "",
          "post_date": "05/26/2022 09:05:44",
          "content": "<p>Thanks. If the only way of tuning is to submit, it's hard.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1801673": "First I want to thank this competition hosts and the Kaggle team for organizing such a interesting competition. And thank you to all the Kagglers.\n\n## Overview\n- Only using SED model\n- Make clean data with [human-in-the-loop](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319813) for weak label\n- Heavy post-processing (public LB 0.79 &rarr; 0.84)\n\nPerhaps my SED is the same model as yours.\n**Clean data** and **post-processing** is important in my solution.\n\n## Clean data (important)\nI made [annotations](https://www.kaggle.com/competitions/birdclef-2021/discussion/239911) in BirdClef2021. And this is effective. But hand labeling is time consuming.\nInstead of hand labeling, I used human-in-the-loop method in this competition.\n- First, train SED with primary label\n- Second, extract high confidence SED predictions in training data\n- I listen these prediction and judge correct or not correct\n- Human verification (I answer yes or no)\n\nI made 2,000 clean data by this human-in-the-loop. These data contain only scored bird.\n\n## Training data\nI used 3 type data.\n- HIL_data (=human in the loop data)\n  - contain [external data](https://www.kaggle.com/code/amandanavine/hawaiian-bird-species)\n- other_data (not scored bird data(about 130species))\n  - initial 3sec audio\n  - The label is primary_label\n- psuedo_data (scored bird, and not contain HIL audio)\n  - initial 3sec audio\n  - psuedo_label = 0.25primay_label + 0.25x1st_generation_model + 0.5x2nd_generation_model \n\nAnd training data is below ratio.\nThis is a best ratio.\n\nHIL_data : other_data : psuedo_data = 1 : 4 : 1\n\n## SED\n- Backbone: eca_nfnet_l0, dm_nfnet_f0\n- Only using \"clipwise_output\" (training & inference)\n- [Basic augment](https://www.kaggle.com/code/shinmurashinmura/birdclef2022-basic-augmentation/notebook)\n  - Time shift\n  - Add pink noise and brown noise\n  - Mix other audio dataset (ESC-50: frog, rain, airplane, crackling_fire)\n- [SpecAugment++](https://www.kaggle.com/competitions/birdclef-2022/discussion/307880) (mixing Mel-spec is ESC-50)\n- Label smoothing (alpha=0.1)\n- Optimize with Adam\n  - lr=0.0001\n  - CosineAnealing (T=10)\n- Epoch:30 (about 14hours in Colaboratory)\n- Loss: BCEWithLogitsLoss\n- Input (train & inference): 5sec\n- STFT resolution:250x254\n  - window_size: 1024\n  - hop_size: 630\n  - mel_bins: 250\n  - fmin: 50\n  - fmax: 14000\n\n## Ensemble\n- Ensemble is a little impact in this competition\n- I compared voting vs average.\n  - Voting is good score a little.\n- Finally, I used voting with 10 models.\n\n## Post-Processing (important)\nMy Post-Processing is similar with [12th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/326979).\n \n#### prediction time shift (=predition_TA)\nprediction_now = now + 0.5next_5sec + 0.25next_10sec + 0.5previous_5sec + 0.25previous_10sec\n#### TTA\nLet t be the target time. I used 3 type inputs.\n- [t,t+5]\n- [t-1,t+4]\n- [t+1,t+6]\n\nAnd these ouput is used in voting.\n#### Threshold optimization (=ThreshO)\n- First, I tuned threhold like [this](https://www.kaggle.com/competitions/birdclef-2022/discussion/318999). These threshold is constant.\n- Second, I tuned threhold each species.\n#### Threshold down (=ThershD)\nIf models detect a certain bird once in the 1 minute audio, I lower the threshold in the audio and infer it once more. \n\n```\nscore = prediction(thresh) # 1min prediction\ndown_species = [\"hawhaw\", \"hawpet1\", \"maupar\", \"ercfra\", \"crehon\", \"puaioh\", \"hawcre\"]\n\nfor bird in down_species:\n    if np.max(score[bird]) > thresh[bird]:\n        thresh[bird] = thresh[bird] - decrease\n\nscore = prediction(thresh) # 1min prediction once more\n```\n\n## Ablation study\n|post-processing|**Public**/Private LB\n|---|---|\n|not using|**0.79**/0.76\n|predition_TA|**0.78**/0.75\n|predition_TA + TTA|**0.78**/0.74\n|predition_TA + TTA + ThreshO|**0.82**/0.77\n|predition_TA + TTA + ThreshO + ThreshD|**0.84**/0.80\n\n## Not working\n- AST (equally SED)\n- PCEN (equally Mel-spec)\n- [ImportantAug](https://www.kaggle.com/competitions/birdclef-2022/discussion/307880)\n- ArcFace as few-shot learning method\n- [STFT Transformer](https://www.kaggle.com/competitions/birdclef-2021/discussion/243360)\n- Cooccurrence species with [these map](https://www.iucnredlist.org/species/22708583/128101101)\n- Using BirdClef2021 dataset",
    "1801703": "Congrats on your solo gold and competition master!\nThank you for sharing your solution. How did you tune threshold for each species?",
    "1801870": "It's trial and trial.\nI ran one line at a time like below.\n\n```\nthresh[label_dic[\"skylar\"]] =  thresh[label_dic[\"skylar\"]] + 0.2\nthresh[label_dic[\"hawcre\"]] =  thresh[label_dic[\"hawcre\"]] + 0.2\nthresh[label_dic[\"aniani\"]] =  thresh[label_dic[\"aniani\"]] + 0.2\nthresh[label_dic[\"hawama\"]] =  thresh[label_dic[\"hawama\"]] + 0.2\n#thresh[label_dic[\"warwhe1\"]] =  thresh[label_dic[\"warwhe1\"]] + 0.2\nthresh[label_dic[\"yefcan\"]] =  thresh[label_dic[\"yefcan\"]] + 0.2\n#thresh[label_dic[\"jabwar\"]] =  thresh[label_dic[\"jabwar\"]] + 0.2\nthresh[label_dic[\"houfin\"]] =  thresh[label_dic[\"houfin\"]] + 0.2\nthresh[label_dic[\"apapan\"]] =  thresh[label_dic[\"apapan\"]] + 0.2\n```\n\nMainly, the threshold for the majority species was raised.",
    "1801903": "Thanks. If the only way of tuning is to submit, it's hard."
  },
  "source": "meta"
}