{
  "id": 243293,
  "title": "4th place solution ",
  "url": "/competitions/birdclef-2021/writeups/third-time-s-the-charm-4th-place-solution",
  "author_name": "",
  "post_date": "2023-09-26T14:04:38.260Z",
  "votes": 85,
  "comment_count": 26,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers of the competition and the Kaggle team for the competition hosting. I would also like to thank all the participants who have shared their knowledge so generously, including the previous Birdcall competitions. Two weeks before the end of this competition, I didn't expect it to be so hard competition.</p>\n<h2>Keys of my solution</h2>\n<ul>\n<li>Ensemble<ul>\n<li>Max 62model (Best private: 47model)</li></ul></li>\n</ul>\n\n<ul>\n<li><strong>Inference with global information for SED model</strong></li>\n<li>Post-processing</li>\n</ul>\n<h2>Training</h2>\n<h3>Preprocess</h3>\n\n<p>The logmelspectrograms were calculated using TorchAudio and normalised using the means and variances per a image.</p>\n<pre><code>logmelspec_extractor = nn.Sequential(\n            MelSpectrogram(\n                ,\n                n_mels=,\n                f_min=,\n                n_fft=,\n                hop_length=,\n                normalized=,\n            ),\n            AmplitudeToDB(top_db=),\n            NormalizeMelSpec(),\n        )\n</code></pre>\n\n<p>For Augmentation in waveform, gaussian and uniform, pink noise are added random.  </p>\n\n<p>As an augmentation on logmelspec, I used Mixup. As a parameter,  I used a probability of 0.2~0.5 and alpha=0.8.</p>\n<h3>Modeling</h3>\n\n<p>I used the SED model like <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">@hidehisaarai1213 's kernel</a>.</p>\n\n<p>Using various backbones, the seconds(max 30s, min 10s) used training has been adjusted so that each has a batchsize of 36.   </p>\n<h3>Training</h3>\n\n<p><code>clipwise_pred</code> is optimized directly. By doing this, I was able to suppress the generation of NaN on backward.</p>\n<pre><code>loss = bce_with_logits(torch.logit(clipwise_pred), target) +  * bce_with_logits((framewise_logit.()[], target)\n</code></pre>\n\n\n<p>Other points are that</p>\n<ul>\n<li>use a secondary label<ul>\n<li>soft labels such as 0.5 did not work</li></ul></li>\n<li>trained 40~50 epochs </li>\n</ul>\n<h3>Pseudo labeling</h3>\n\n<p>I have tried some patterns using a combination of the following steps.</p>\n\n<ul>\n<li>Simple per-audio file relabeling<ul>\n<li>Use clipwise_pred and threshold = 0.2</li></ul></li>\n</ul>\n\n<ul>\n<li>Inspired by <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183204\" target=\"_blank\">@hidehisaarai1213 's solution</a> in the last competition, I saved <code>framewise_pred</code> and <code>time_att</code> with the whole audio file as input, and computed <code>clipwise_pred</code> by cropping the period used for training.</li>\n</ul>\n\n<ul>\n<li>Create a pseudo label using the second method only for audio files without a secondary_label</li>\n</ul>\n\n<ul>\n<li>If the pseudo label was below a certain value (0.05), the probability of 0.1 was used to set the primary and secondary labels to 0.</li>\n</ul>\n\n<p>All patterns improved the performance of the single models by about 0.01, and although the score on train_soundscape did not contribute to ensemble, it did work somewhat on the public LB.</p>\n<h3>30s finetuning</h3>\n\n<p>Larger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs. There is a slight improvement in score with train_soundscape, and it is included in ensemble.</p>\n<h3>Checkpoint selection</h3>\n\n<p>The checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.</p>\n<h2>Inference with global information</h2>\n\n<p>This is my favorite and most effective part of my solution.  </p>\n\n<p>To begin with, the idea that I had before joining this competition was that if the SED model could accurately perform \"Sound Event Detection\", then it would be possible to improve predictions for the shorter time segment by cropping from the longer time segment the necessary parts of predictions.</p>\n\n<p>In this way, longer term features can be used for inference.</p>\n<pre><code>framewise_pred_5s = self.fix_scale(feat[:, :, start:end])\natt_5s = torch.softmax(time_att[:, :, start:end], dim=-)\nclipwise_pred_5s = torch.(torch.sigmoid(framewise_pred_5s) * att_5s, dim=-,)\n</code></pre>\n\n<p>By implementing this idea, I can improve both score and public LB in train_soundscape by about 0.03. In practice, I used 30s as the segment for the longer period and implemented it so that the 5s interval I want to predict is in the center. This idea is also used to create a pseudo label in the second pattern.  </p>\n\n<p>Secondly, in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">7th place solution in previous Birdcall competition</a>,  the birds that were observed to appear using a high threshold for each audio file can be expected to appear in short segments with a high probability, and the sensitivity can be increased by lowering the threshold.  </p>\n\n<p>Inspired by this observation, I used a high threshold (0.05) in <code>clipwise_pred_30s</code> to generate a list of possible birds and a low threshold (0.025) in <code>clipwise_pred_5s</code>,  to perform AND operations.</p>\n<pre><code>((clipwise_pred_30s &gt; high_threshold) + (clipwise_pred_5s &gt; low_threshold)) &gt;= \n</code></pre>\n\n<p>To get an idea of these thresholds, I used <code>scipy.optimize.dual_annealing</code> to optimize for train_soundscape. This seems rather risky, but I thought it was somewhat reasonable due to the simple public LB probing described below.</p>\n\n<p>This double thresholding improved the score of public LB and train_soundscape by 0.03.  </p>\n\n<p>In the ensemble I took a simple average. Initially, hard voting was considered, but the score did not change much, so the simple method was chosen.  </p>\n\n<h2>Post-processing with location and date</h2>\n\n<p>A list of birds observed within 450 meters around each site and a list of birds appearing in each month was made, and those not appearing were removed from the submission.</p>\n\n<p>Here, because the two sites \"COR\" and \"COL\" are relatively close, the birds have only been removed if both are zero.</p>\n\n<p>I have also quadrupled the thresholds for some rare classes of birds observed in the vicinity of the sites.</p>\n\n<p>This post-processing resulted in a consistent improvement of less than 0.01 for both train_soundscape and public LB.</p>\n<h2>Simple public LB probing</h2>\n\n<p>Since train_soundscape and public LB were correlated to some extent, I was concerned that the distribution of public LB might be extremely similar to train_soundscape.</p>\n\n<p>To verify this, I changed the site and bird predictions not included in the train_soundscape to nobird.</p>\n\n<p>If there is no change in the score, it is assumed that train_soundscape and public LB have the same distribution, but the score has decreased considerably. (bird: 0.71-&gt;0.64, site: 0.71-&gt;0.58)</p>\n\n<p>So I can see that public LB contains sites and birds that are not included in train_soundscape.</p>\n\n<p>From this, I forcefully predicted that public and private would be randomly split. (This is strictly uncertain, but there was nothing else I could do.)</p>\n<h2>Experimental code and the inference notebook for best submission</h2>\n<p>Code: <a href=\"https://github.com/tattaka/birdclef-2021\" target=\"_blank\">https://github.com/tattaka/birdclef-2021</a><br>\nInference notebook: <a href=\"https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\" target=\"_blank\">https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465</a><br>\nAggregation of the number of birds for post-processing: <a href=\"https://www.kaggle.com/tattaka/make-month-and-site-mask\" target=\"_blank\">https://www.kaggle.com/tattaka/make-month-and-site-mask</a></p>",
  "messages": [
    {
      "id": "1332153",
      "postDate": "06/02/2021 00:18:56",
      "content": "<p>First of all, I would like to thank the organizers of the competition and the Kaggle team for the competition hosting. I would also like to thank all the participants who have shared their knowledge so generously, including the previous Birdcall competitions. Two weeks before the end of this competition, I didn't expect it to be so hard competition.</p>\n<h2>Keys of my solution</h2>\n<ul>\n<li>Ensemble<ul>\n<li>Max 62model (Best private: 47model)</li></ul></li>\n</ul>\n\n<ul>\n<li><strong>Inference with global information for SED model</strong></li>\n<li>Post-processing</li>\n</ul>\n<h2>Training</h2>\n<h3>Preprocess</h3>\n\n<p>The logmelspectrograms were calculated using TorchAudio and normalised using the means and variances per a image.</p>\n<pre><code>logmelspec_extractor = nn.Sequential(\n            MelSpectrogram(\n                ,\n                n_mels=,\n                f_min=,\n                n_fft=,\n                hop_length=,\n                normalized=,\n            ),\n            AmplitudeToDB(top_db=),\n            NormalizeMelSpec(),\n        )\n</code></pre>\n\n<p>For Augmentation in waveform, gaussian and uniform, pink noise are added random.  </p>\n\n<p>As an augmentation on logmelspec, I used Mixup. As a parameter,  I used a probability of 0.2~0.5 and alpha=0.8.</p>\n<h3>Modeling</h3>\n\n<p>I used the SED model like <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">@hidehisaarai1213 's kernel</a>.</p>\n\n<p>Using various backbones, the seconds(max 30s, min 10s) used training has been adjusted so that each has a batchsize of 36.   </p>\n<h3>Training</h3>\n\n<p><code>clipwise_pred</code> is optimized directly. By doing this, I was able to suppress the generation of NaN on backward.</p>\n<pre><code>loss = bce_with_logits(torch.logit(clipwise_pred), target) +  * bce_with_logits((framewise_logit.()[], target)\n</code></pre>\n\n\n<p>Other points are that</p>\n<ul>\n<li>use a secondary label<ul>\n<li>soft labels such as 0.5 did not work</li></ul></li>\n<li>trained 40~50 epochs </li>\n</ul>\n<h3>Pseudo labeling</h3>\n\n<p>I have tried some patterns using a combination of the following steps.</p>\n\n<ul>\n<li>Simple per-audio file relabeling<ul>\n<li>Use clipwise_pred and threshold = 0.2</li></ul></li>\n</ul>\n\n<ul>\n<li>Inspired by <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183204\" target=\"_blank\">@hidehisaarai1213 's solution</a> in the last competition, I saved <code>framewise_pred</code> and <code>time_att</code> with the whole audio file as input, and computed <code>clipwise_pred</code> by cropping the period used for training.</li>\n</ul>\n\n<ul>\n<li>Create a pseudo label using the second method only for audio files without a secondary_label</li>\n</ul>\n\n<ul>\n<li>If the pseudo label was below a certain value (0.05), the probability of 0.1 was used to set the primary and secondary labels to 0.</li>\n</ul>\n\n<p>All patterns improved the performance of the single models by about 0.01, and although the score on train_soundscape did not contribute to ensemble, it did work somewhat on the public LB.</p>\n<h3>30s finetuning</h3>\n\n<p>Larger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs. There is a slight improvement in score with train_soundscape, and it is included in ensemble.</p>\n<h3>Checkpoint selection</h3>\n\n<p>The checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.</p>\n<h2>Inference with global information</h2>\n\n<p>This is my favorite and most effective part of my solution.  </p>\n\n<p>To begin with, the idea that I had before joining this competition was that if the SED model could accurately perform \"Sound Event Detection\", then it would be possible to improve predictions for the shorter time segment by cropping from the longer time segment the necessary parts of predictions.</p>\n\n<p>In this way, longer term features can be used for inference.</p>\n<pre><code>framewise_pred_5s = self.fix_scale(feat[:, :, start:end])\natt_5s = torch.softmax(time_att[:, :, start:end], dim=-)\nclipwise_pred_5s = torch.(torch.sigmoid(framewise_pred_5s) * att_5s, dim=-,)\n</code></pre>\n\n<p>By implementing this idea, I can improve both score and public LB in train_soundscape by about 0.03. In practice, I used 30s as the segment for the longer period and implemented it so that the 5s interval I want to predict is in the center. This idea is also used to create a pseudo label in the second pattern.  </p>\n\n<p>Secondly, in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">7th place solution in previous Birdcall competition</a>,  the birds that were observed to appear using a high threshold for each audio file can be expected to appear in short segments with a high probability, and the sensitivity can be increased by lowering the threshold.  </p>\n\n<p>Inspired by this observation, I used a high threshold (0.05) in <code>clipwise_pred_30s</code> to generate a list of possible birds and a low threshold (0.025) in <code>clipwise_pred_5s</code>,  to perform AND operations.</p>\n<pre><code>((clipwise_pred_30s &gt; high_threshold) + (clipwise_pred_5s &gt; low_threshold)) &gt;= \n</code></pre>\n\n<p>To get an idea of these thresholds, I used <code>scipy.optimize.dual_annealing</code> to optimize for train_soundscape. This seems rather risky, but I thought it was somewhat reasonable due to the simple public LB probing described below.</p>\n\n<p>This double thresholding improved the score of public LB and train_soundscape by 0.03.  </p>\n\n<p>In the ensemble I took a simple average. Initially, hard voting was considered, but the score did not change much, so the simple method was chosen.  </p>\n\n<h2>Post-processing with location and date</h2>\n\n<p>A list of birds observed within 450 meters around each site and a list of birds appearing in each month was made, and those not appearing were removed from the submission.</p>\n\n<p>Here, because the two sites \"COR\" and \"COL\" are relatively close, the birds have only been removed if both are zero.</p>\n\n<p>I have also quadrupled the thresholds for some rare classes of birds observed in the vicinity of the sites.</p>\n\n<p>This post-processing resulted in a consistent improvement of less than 0.01 for both train_soundscape and public LB.</p>\n<h2>Simple public LB probing</h2>\n\n<p>Since train_soundscape and public LB were correlated to some extent, I was concerned that the distribution of public LB might be extremely similar to train_soundscape.</p>\n\n<p>To verify this, I changed the site and bird predictions not included in the train_soundscape to nobird.</p>\n\n<p>If there is no change in the score, it is assumed that train_soundscape and public LB have the same distribution, but the score has decreased considerably. (bird: 0.71-&gt;0.64, site: 0.71-&gt;0.58)</p>\n\n<p>So I can see that public LB contains sites and birds that are not included in train_soundscape.</p>\n\n<p>From this, I forcefully predicted that public and private would be randomly split. (This is strictly uncertain, but there was nothing else I could do.)</p>\n<h2>Experimental code and the inference notebook for best submission</h2>\n<p>Code: <a href=\"https://github.com/tattaka/birdclef-2021\" target=\"_blank\">https://github.com/tattaka/birdclef-2021</a><br>\nInference notebook: <a href=\"https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\" target=\"_blank\">https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465</a><br>\nAggregation of the number of birds for post-processing: <a href=\"https://www.kaggle.com/tattaka/make-month-and-site-mask\" target=\"_blank\">https://www.kaggle.com/tattaka/make-month-and-site-mask</a></p>",
      "rawMarkdown": "First of all, I would like to thank the organizers of the competition and the Kaggle team for the competition hosting. I would also like to thank all the participants who have shared their knowledge so generously, including the previous Birdcall competitions. Two weeks before the end of this competition, I didn't expect it to be so hard competition.\n\n## Keys of my solution\n* Ensemble\n  * Max 62model (Best private: 47model)\n<!-- * global情報を用いたinference -->\n* **Inference with global information for SED model**\n* Post-processing\n\n## Training\n### Preprocess\n<!-- logmelspectrogramはTorchAudioを用いて計算され、それぞれの平均と分散を用いて正規化されました。 -->\nThe logmelspectrograms were calculated using TorchAudio and normalised using the means and variances per a image.\n``` python\nlogmelspec_extractor = nn.Sequential(\n            MelSpectrogram(\n                32000,\n                n_mels=128,\n                f_min=20,\n                n_fft=2048,\n                hop_length=512,\n                normalized=True,\n            ),\n            AmplitudeToDB(top_db=80.0),\n            NormalizeMelSpec(),\n        )\n```  \n<!-- waveformでのAugmentationとして、pink noiseやgaussian noiseを加えました。 -->\nFor Augmentation in waveform, gaussian and uniform, pink noise are added random.  \n<!-- logmelspec上のaugmentationとして、Mixupを用いました。パラメータとして0.2~0.5の確率、alpha=0.8を用いました。 -->\nAs an augmentation on logmelspec, I used Mixup. As a parameter,  I used a probability of 0.2~0.5 and alpha=0.8.\n### Modeling\n<!-- araiさんのkernelと同じくSED modelを用いました。 -->\nI used the SED model like [@hidehisaarai1213 's kernel](https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter).\n<!-- 様々なbackboneを用い、それぞれbatchsizeが36になるように使用する秒数が調整されました(最大で30s, 最小で10s) -->\nUsing various backbones, the seconds(max 30s, min 10s) used training has been adjusted so that each has a batchsize of 36.   \n### Training\n<!-- @hidehisaarai1213のmodelとはloss関数の計算が少し異なり`clipwise_pred`を直接最適化しています。こうすることでbackward時のNaNの発生を抑制することができました。 -->\n`clipwise_pred` is optimized directly. By doing this, I was able to suppress the generation of NaN on backward.\n``` python\nloss = bce_with_logits(torch.logit(clipwise_pred), target) + 0.5 * bce_with_logits((framewise_logit.max(1)[0], target)\n```\n<!-- loss関数はnn.BCEWithLogitsLossを用いました。 -->\n<!-- The loss function used is nn.BCEWithLogitsLoss. -->\n\nOther points are that\n* use a secondary label\n  * soft labels such as 0.5 did not work\n* trained 40~50 epochs \n  \n### Pseudo labeling\n<!-- 私は以下の手順を組み合わせていくつかのパターンを試してみました。 -->\nI have tried some patterns using a combination of the following steps.\n<!-- * 単純にaudio file単位のrelabeling\n  * clipwise_predを用い、閾値は0.2を用いました -->\n* Simple per-audio file relabeling\n  * Use clipwise_pred and threshold = 0.2\n<!-- * 前回のコンペのaraiさんのsolutionにインスピレーションを受け、audio file全体を入力としたframewise_predとtime_attを保存し、学習に用いる期間分を切り取りclipwise_predを計算しました。 -->\n* Inspired by [@hidehisaarai1213 's solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183204) in the last competition, I saved `framewise_pred` and `time_att` with the whole audio file as input, and computed `clipwise_pred` by cropping the period used for training.\n<!-- * secondary_labelがないaudio fileだけ2番目の方法でpseudo labelを作成する -->\n* Create a pseudo label using the second method only for audio files without a secondary_label\n<!-- * 閾値がある一定より下(0.05)だった場合、0.1の確率でprimary labelとsecondary labelを削除しました。 -->\n* If the pseudo label was below a certain value (0.05), the probability of 0.1 was used to set the primary and secondary labels to 0.\n\n<!-- どのパターンも個々のモデルの性能は0.01ほど向上し、train_soundscapeでのscoreはensembleに貢献しませんでしたが、public LB上では多少機能しました。 -->\nAll patterns improved the performance of the single models by about 0.01, and although the score on train_soundscape did not contribute to ensemble, it did work somewhat on the public LB.\n\n### 30s finetuning\n<!-- 大きいモデルではbatchsizeを確保するために10sなどの短いsegmentで学習されています。そのため、ラベルがノイジーになります。これを解決するためにAttention moduleのみを30sのsegmentを用いて10epoch追加で学習させました。train_soundscapeでのscoreでわずかな改善があり、ensembleの中に含まれています。 -->\nLarger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs. There is a slight improvement in score with train_soundscape, and it is included in ensemble.\n\n### Checkpoint selection\n<!-- それぞれvalidationでもtrainingと同じ秒数で切り取ったsegmentを用いout-of-foldでのlossが最も小さいものを選びました。 -->\nThe checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.\n\n## Inference with global information\n<!-- 自分のsolutionのなかで最も効果的で気に入っている部分です。   -->\nThis is my favorite and most effective part of my solution.  \n<!-- はじめに、このcompetitionに参加する前から持っていたアイディアは、SEDモデルが正確に\"Sound Event Detection\"を実行できるのであれば、長い期間に対して行われた予測の中で必要な部分だけを切り取ることで短い期間の予測を改善することができるのではないかということでした。 -->\nTo begin with, the idea that I had before joining this competition was that if the SED model could accurately perform \"Sound Event Detection\", then it would be possible to improve predictions for the shorter time segment by cropping from the longer time segment the necessary parts of predictions.\n\n\n<!-- このようにすることで、より長期間の特徴を推測に用いることができます。 -->\nIn this way, longer term features can be used for inference.\n```python\nframewise_pred_5s = self.fix_scale(feat[:, :, start:end])\natt_5s = torch.softmax(time_att[:, :, start:end], dim=-1)\nclipwise_pred_5s = torch.sum(torch.sigmoid(framewise_pred_5s) * att_5s, dim=-1,)\n```\n<!-- このアイディアを実行することで0.03ほどtrain_soundscapeでのscoreとpublic LB共に改善することができます。実際には長い期間のsegmentとして30sを用い、予測したい5sの区間が中心にくるように実装しました。このアイディアはpseudo labelを作る2つめのパターンでも用いられています。   -->\nBy implementing this idea, I can improve both score and public LB in train_soundscape by about 0.03. In practice, I used 30s as the segment for the longer period and implemented it so that the 5s interval I want to predict is in the center. This idea is also used to create a pseudo label in the second pattern.  \n<!-- 次に、前回の[Birdcallでの7th解法](https://www.kaggle.com/c/birdsong-recognition/discussion/183571)での観察ではaudio file単位で高い閾値を使い出現することを確認できた鳥は、短いsegmentにおいても高い確率で出現することが期待でき、閾値を下げて感度を上げることができます。 -->\nSecondly, in [7th place solution in previous Birdcall competition](https://www.kaggle.com/c/birdsong-recognition/discussion/183571),  the birds that were observed to appear using a high threshold for each audio file can be expected to appear in short segments with a high probability, and the sensitivity can be increased by lowering the threshold.  \n<!-- この観察に触発され、`clipwise_pred_30s`では高い閾値(0.05)を用いて考えられる鳥のリストを作成し、`clipwise_pred_5s`では低い閾値(0.025)を用いAND演算を行いました。 -->\nInspired by this observation, I used a high threshold (0.05) in `clipwise_pred_30s` to generate a list of possible birds and a low threshold (0.025) in `clipwise_pred_5s`,  to perform AND operations.\n```python\n((clipwise_pred_30s > high_threshold) + (clipwise_pred_5s > low_threshold)) >= 2\n```\n<!-- これらの閾値の見当をつけるために`scipy.optimize.dual_annealing`を用い、train_soundscapeに対して最適化しました。これはかなり危険なように見えますが、後に述べる単純なpublic LB probeによりある程度妥当だと自分は考えました。 -->\nTo get an idea of these thresholds, I used `scipy.optimize.dual_annealing` to optimize for train_soundscape. This seems rather risky, but I thought it was somewhat reasonable due to the simple public LB probing described below.\n<!-- この二重の閾値処理によってpublic LBとtrain_soundscapeのscoreは0.03改善されました。   -->\nThis double thresholding improved the score of public LB and train_soundscape by 0.03.  \n<!-- アンサンブルでは単純な平均を取りました。初期にhard voteも検討しましたが、あまりスコアは変わらないため、単純な手法を選びました。   -->\nIn the ensemble I took a simple average. Initially, hard voting was considered, but the score did not change much, so the simple method was chosen.  \n<!-- ## 場所と日付を用いたpostprocessing -->\n## Post-processing with location and date\n<!-- それぞれのsite周辺450メートル内で観測された鳥のリストと、それぞれの月ごとに出現する鳥のリストを作成し、出現しないものをsubmissionから削除しました。 -->\nA list of birds observed within 450 meters around each site and a list of birds appearing in each month was made, and those not appearing were removed from the submission.\n<!-- ただし、\"COR\"と\"COL\"の2つのsiteは比較的近くのため、どちらも0の場合のみ削除されました。 -->\nHere, because the two sites \"COR\" and \"COL\" are relatively close, the birds have only been removed if both are zero.\n<!-- また、site周辺で観察された鳥の中でいくつかの希少なクラスに関して閾値を4倍にしました。   -->\nI have also quadrupled the thresholds for some rare classes of birds observed in the vicinity of the sites.\n<!-- これらのpostprocessingにより、train_soundscape・public LB共に0.01以下の一貫した改善が見られました。 -->\nThis post-processing resulted in a consistent improvement of less than 0.01 for both train_soundscape and public LB.\n\n## Simple public LB probing\n<!-- ある程度train_soundscapeとpublic LBは相関していたため、public LBの分布が極端にtrain_soundscapeと似ているのではないかという危惧を感じました。 -->\nSince train_soundscape and public LB were correlated to some extent, I was concerned that the distribution of public LB might be extremely similar to train_soundscape.\n<!-- それを検証するために、適当なsubmissionに対して、train_soundscapeに含まれていないsiteと鳥の予測をnobirdに変えてsubmitしました。 -->\nTo verify this, I changed the site and bird predictions not included in the train_soundscape to nobird.\n<!-- ここで変化がなければtrain_soundscapeととpublic LBは同じ分布であると考えられますが、かなりscoreが低下しました。(0.71->0.58) -->\nIf there is no change in the score, it is assumed that train_soundscape and public LB have the same distribution, but the score has decreased considerably. (bird: 0.71->0.64, site: 0.71->0.58)\n<!-- よってpublic LBにはtrain_soundscapeに含まれないsite・鳥が含まれることがわかります。 -->\nSo I can see that public LB contains sites and birds that are not included in train_soundscape.\n<!-- ここから、強引な考えですが、publicとprivate はランダムに分割されていると予測しました。(これは厳密には不確実ではありますが、他に手のうちようがありませんでした) -->\nFrom this, I forcefully predicted that public and private would be randomly split. (This is strictly uncertain, but there was nothing else I could do.)\n\n## Experimental code and the inference notebook for best submission\nCode: https://github.com/tattaka/birdclef-2021\nInference notebook: https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\nAggregation of the number of birds for post-processing: https://www.kaggle.com/tattaka/make-month-and-site-mask",
      "votes": null
    },
    {
      "id": "1332155",
      "postDate": "06/02/2021 00:22:15",
      "content": "<p>Congrats for the solo gold and thank you for sharing your unique approach. Does the post-processing also work for private lb?</p>",
      "rawMarkdown": "Congrats for the solo gold and thank you for sharing your unique approach. Does the post-processing also work for private lb?",
      "votes": null
    },
    {
      "id": "1332157",
      "postDate": "06/02/2021 00:24:45",
      "content": "<p>Congratulations 🎉 <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> we also did a similar post processing. Good job with \"Inference from global information\" 👋</p>",
      "rawMarkdown": "Congratulations 🎉 @tattaka we also did a similar post processing. Good job with \"Inference from global information\" 👋",
      "votes": null
    },
    {
      "id": "1332159",
      "postDate": "06/02/2021 00:25:51",
      "content": "<p>It worked for us 😊</p>",
      "rawMarkdown": "It worked for us 😊",
      "votes": null
    },
    {
      "id": "1332161",
      "postDate": "06/02/2021 00:26:26",
      "content": "<p>Seems to work the same<br>\n<img src=\"https://user-images.githubusercontent.com/16153860/120405827-6c007200-c384-11eb-86b7-456b2e686e9e.png\" alt=\"\"></p>",
      "rawMarkdown": "Seems to work the same\n![](https://user-images.githubusercontent.com/16153860/120405827-6c007200-c384-11eb-86b7-456b2e686e9e.png)",
      "votes": null
    },
    {
      "id": "1332165",
      "postDate": "06/02/2021 00:28:21",
      "content": "<p>Congratulations! 🎉🎉🎉🎉<br>\nWe also tried  the same post-processing by site location, but I used 80km, 200km, and 500km. Maybe this is the reason why I have no effect</p>",
      "rawMarkdown": "Congratulations! 🎉🎉🎉🎉\nWe also tried  the same post-processing by site location, but I used 80km, 200km, and 500km. Maybe this is the reason why I have no effect",
      "votes": null
    },
    {
      "id": "1332168",
      "postDate": "06/02/2021 00:29:10",
      "content": "<p>Congrats on the solo gold and the 4th place!</p>\n<p>Thank you for sharing the interesting approach.</p>",
      "rawMarkdown": "Congrats on the solo gold and the 4th place!\n\nThank you for sharing the interesting approach.",
      "votes": null
    },
    {
      "id": "1332178",
      "postDate": "06/02/2021 00:40:57",
      "content": "<p>Thank you for the info for both of you.</p>",
      "rawMarkdown": "Thank you for the info for both of you.",
      "votes": null
    },
    {
      "id": "1332180",
      "postDate": "06/02/2021 00:43:11",
      "content": "<p>Congratulations!<br>\nCould you please show me the code for the NormalizeMelSpec() function?</p>",
      "rawMarkdown": "Congratulations!\nCould you please show me the code for the NormalizeMelSpec() function?",
      "votes": null
    },
    {
      "id": "1332181",
      "postDate": "06/02/2021 00:44:29",
      "content": "<p>Big Congratulations on your solo gold…<br>\nIf you dont mind can you please share how long did it take for you train the SED model for 50 epochs and the hardware you used??</p>",
      "rawMarkdown": "Big Congratulations on your solo gold...\nIf you dont mind can you please share how long did it take for you train the SED model for 50 epochs and the hardware you used??",
      "votes": null
    },
    {
      "id": "1332184",
      "postDate": "06/02/2021 00:50:49",
      "content": "<p>That's actually a great question, i tried to train a SED model in colab pro and they banned me for abuse</p>",
      "rawMarkdown": "That's actually a great question, i tried to train a SED model in colab pro and they banned me for abuse",
      "votes": null
    },
    {
      "id": "1332186",
      "postDate": "06/02/2021 00:51:48",
      "content": "<p>Some of the mono_to_color functions used in the previous Birdcall competition have been applied to (bs, 1, mel, time).</p>\n<pre><code>class NormalizeMelSpec(nn.Module):\n    def __init__(self, eps=1e-6):\n        super().__init__()\n        self.eps = eps\n\n    def forward(self, X):\n        mean = X.mean((1, 2), keepdim=True)\n        std = X.std((1, 2), keepdim=True)\n        Xstd = (X - mean) / (std + self.eps)\n        norm_min, norm_max = Xstd.min(-1)[0].min(-1)[0], Xstd.max(-1)[0].max(-1)[0]\n        fix_ind = (norm_max - norm_min) &gt; self.eps * torch.ones_like(\n            (norm_max - norm_min)\n        )\n        V = torch.zeros_like(Xstd)\n        if fix_ind.sum():\n            V_fix = Xstd[fix_ind]\n            norm_max_fix = norm_max[fix_ind, None, None]\n            norm_min_fix = norm_min[fix_ind, None, None]\n            V_fix = torch.max(\n                torch.min(V_fix, norm_max_fix),\n                norm_min_fix,\n            )\n            # print(V_fix.shape, norm_min_fix.shape, norm_max_fix.shape)\n            V_fix = (V_fix - norm_min_fix) / (norm_max_fix - norm_min_fix)\n            V[fix_ind] = V_fix\n        return V\n</code></pre>",
      "rawMarkdown": "Some of the mono_to_color functions used in the previous Birdcall competition have been applied to (bs, 1, mel, time).\n``` python\nclass NormalizeMelSpec(nn.Module):\n    def __init__(self, eps=1e-6):\n        super().__init__()\n        self.eps = eps\n\n    def forward(self, X):\n        mean = X.mean((1, 2), keepdim=True)\n        std = X.std((1, 2), keepdim=True)\n        Xstd = (X - mean) / (std + self.eps)\n        norm_min, norm_max = Xstd.min(-1)[0].min(-1)[0], Xstd.max(-1)[0].max(-1)[0]\n        fix_ind = (norm_max - norm_min) > self.eps * torch.ones_like(\n            (norm_max - norm_min)\n        )\n        V = torch.zeros_like(Xstd)\n        if fix_ind.sum():\n            V_fix = Xstd[fix_ind]\n            norm_max_fix = norm_max[fix_ind, None, None]\n            norm_min_fix = norm_min[fix_ind, None, None]\n            V_fix = torch.max(\n                torch.min(V_fix, norm_max_fix),\n                norm_min_fix,\n            )\n            # print(V_fix.shape, norm_min_fix.shape, norm_max_fix.shape)\n            V_fix = (V_fix - norm_min_fix) / (norm_max_fix - norm_min_fix)\n            V[fix_ind] = V_fix\n        return V\n\n```",
      "votes": null
    },
    {
      "id": "1332189",
      "postDate": "06/02/2021 00:52:51",
      "content": "<p>I used 1080ti x 3, it takes about 10~15 hours to train 50epoch.</p>",
      "rawMarkdown": "I used 1080ti x 3, it takes about 10~15 hours to train 50epoch.",
      "votes": null
    },
    {
      "id": "1332193",
      "postDate": "06/02/2021 00:55:55",
      "content": "<p>I was stalling to implement a location based post processing for a long time, just the thought about dealing with name files on the private dataset was enough to make me focus on something else haha.</p>\n<p>Now that i know that it works i kind of want to implement it just to see the impact</p>",
      "rawMarkdown": "I was stalling to implement a location based post processing for a long time, just the thought about dealing with name files on the private dataset was enough to make me focus on something else haha.\n\nNow that i know that it works i kind of want to implement it just to see the impact",
      "votes": null
    },
    {
      "id": "1332194",
      "postDate": "06/02/2021 00:57:02",
      "content": "<p>That's not much actually! Congratz for the clean code!</p>",
      "rawMarkdown": "That's not much actually! Congratz for the clean code!",
      "votes": null
    },
    {
      "id": "1332211",
      "postDate": "06/02/2021 01:22:15",
      "content": "<p>Thank you. I'll try.</p>",
      "rawMarkdown": "Thank you. I'll try.",
      "votes": null
    },
    {
      "id": "1332406",
      "postDate": "06/02/2021 05:41:02",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> on 4th place and thanks for sharing solution </p>",
      "rawMarkdown": "Congrats @tattaka on 4th place and thanks for sharing solution",
      "votes": null
    },
    {
      "id": "1332689",
      "postDate": "06/02/2021 08:42:36",
      "content": "<p>Congratulations , i've one question, did you use spec augmentations ? </p>",
      "rawMarkdown": "Congratulations , i've one question, did you use spec augmentations ?",
      "votes": null
    },
    {
      "id": "1332711",
      "postDate": "06/02/2021 08:54:46",
      "content": "<p>No, it didn't work for my solution.</p>",
      "rawMarkdown": "No, it didn't work for my solution.",
      "votes": null
    },
    {
      "id": "1332730",
      "postDate": "06/02/2021 09:08:23",
      "content": "<p>Congrats to the score and great solution! 62 models, that’s an ensemble to remember!</p>",
      "rawMarkdown": "Congrats to the score and great solution! 62 models, that’s an ensemble to remember!",
      "votes": null
    },
    {
      "id": "1332747",
      "postDate": "06/02/2021 09:18:17",
      "content": "<p>Congrats on the result.  I'll reread your writeup as I am now convinced that training on 5 seconds clips as I did is not the way to go.  You made SED models work quite well.</p>",
      "rawMarkdown": "Congrats on the result.  I'll reread your writeup as I am now convinced that training on 5 seconds clips as I did is not the way to go.  You made SED models work quite well.",
      "votes": null
    },
    {
      "id": "1334449",
      "postDate": "06/03/2021 14:01:56",
      "content": "<p>I have released the experimental code and the inference notebook for the best submission.</p>\n<p>Code: <a href=\"https://github.com/tattaka/birdclef-2021\" target=\"_blank\">https://github.com/tattaka/birdclef-2021</a><br>\nInference notebook: <a href=\"https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\" target=\"_blank\">https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465</a><br>\nAggregation of the number of birds for post-processing: <a href=\"https://www.kaggle.com/tattaka/make-month-and-site-mask\" target=\"_blank\">https://www.kaggle.com/tattaka/make-month-and-site-mask</a></p>\n<p>In cleaning the code, I've removed duplicate experiments and non-reproducible ones that didn't specify a seed, but you should get similar scores.<br>\nPlease forgive me if my inference notebook is not organized due to laziness and time-saving.</p>",
      "rawMarkdown": "I have released the experimental code and the inference notebook for the best submission.\n\nCode: https://github.com/tattaka/birdclef-2021\nInference notebook: https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\nAggregation of the number of birds for post-processing: https://www.kaggle.com/tattaka/make-month-and-site-mask\n\nIn cleaning the code, I've removed duplicate experiments and non-reproducible ones that didn't specify a seed, but you should get similar scores.\nPlease forgive me if my inference notebook is not organized due to laziness and time-saving.",
      "votes": null
    },
    {
      "id": "1346879",
      "postDate": "06/12/2021 18:11:17",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> on 4th place and thanks for sharing your code and notebooks! <br>\nI was wondering about your thresholds. Do you have an explanation why your optimal threshold (around 0.05) is so much lower, compared to other solutions (usually around 0.3 even for other SED approaches)?</p>",
      "rawMarkdown": "Congratulations @tattaka on 4th place and thanks for sharing your code and notebooks! \nI was wondering about your thresholds. Do you have an explanation why your optimal threshold (around 0.05) is so much lower, compared to other solutions (usually around 0.3 even for other SED approaches)?",
      "votes": null
    },
    {
      "id": "1348410",
      "postDate": "06/14/2021 03:26:33",
      "content": "<p>Thank you for sharing!</p>\n<p>I have a question about \"Inference with global information\".<br>\nI checked your code. And my recognition is </p>\n<ul>\n<li>get long clip feat (eg. 10sec)</li>\n<li>get feat[0sec:5sec] and feat[5sec:10sec]</li>\n<li>get clipwise_prediction(feat[0:5]) and clipwise_prediction(feat[5:10])</li>\n</ul>\n<p>The effect of your idea is that we can get information about <strong>the seams</strong> of the clip.<br>\nIs this this recognition correct?</p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/6258610a-d947-fb75-acb3-981b9db206c0.png\" alt=\"image.png\"></p>",
      "rawMarkdown": "Thank you for sharing!\n\nI have a question about \"Inference with global information\".\nI checked your code. And my recognition is \n+ get long clip feat (eg. 10sec)\n+ get feat[0sec:5sec] and feat[5sec:10sec]\n+ get clipwise_prediction(feat[0:5]) and clipwise_prediction(feat[5:10])\n\nThe effect of your idea is that we can get information about **the seams** of the clip.\nIs this this recognition correct?\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/6258610a-d947-fb75-acb3-981b9db206c0.png)",
      "votes": null
    },
    {
      "id": "1350497",
      "postDate": "06/15/2021 13:35:06",
      "content": "<p>With standard 5s prediction, the optimal threshold was close to 0.2.<br>\nMaybe, the use of global and local prediction can be used to stabilize and increase the sensitivity of your inferences.</p>",
      "rawMarkdown": "With standard 5s prediction, the optimal threshold was close to 0.2.\nMaybe, the use of global and local prediction can be used to stabilize and increase the sensitivity of your inferences.",
      "votes": null
    },
    {
      "id": "1350519",
      "postDate": "06/15/2021 14:01:10",
      "content": "<p>Instead of getting two predictions from 10s to 5s, I get predictions from the central 5s and the whole 10s. Therefore, the number of inferences is the same as the normal 5s prediction.<br>\nThe disadvantage is that it takes longer to create a logmelspec than with 5s.   <br>\nAlso, inference near the edge of the clip repeats the clip to make up for the missing.<br>\nSo the global inference at the edge of the clip may be incorrect: (<br>\n<img src=\"https://user-images.githubusercontent.com/16153860/122066104-68adc180-ce2d-11eb-98b1-8869ba7cd4bd.png\" alt=\"\"></p>",
      "rawMarkdown": "Instead of getting two predictions from 10s to 5s, I get predictions from the central 5s and the whole 10s. Therefore, the number of inferences is the same as the normal 5s prediction.\nThe disadvantage is that it takes longer to create a logmelspec than with 5s.   \nAlso, inference near the edge of the clip repeats the clip to make up for the missing.\nSo the global inference at the edge of the clip may be incorrect: (\n![](https://user-images.githubusercontent.com/16153860/122066104-68adc180-ce2d-11eb-98b1-8869ba7cd4bd.png)",
      "votes": null
    },
    {
      "id": "1351033",
      "postDate": "06/16/2021 03:03:26",
      "content": "<p>Thanks for the reply.</p>\n<p>I had misunderstood.<br>\nBut your diagram helped me.</p>",
      "rawMarkdown": "Thanks for the reply.\n\nI had misunderstood.\nBut your diagram helped me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1332155,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/02/2021 00:22:15",
      "content": "<p>Congrats for the solo gold and thank you for sharing your unique approach. Does the post-processing also work for private lb?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332159,
          "author_name": "lplenka",
          "author_url": "",
          "post_date": "06/02/2021 00:25:51",
          "content": "<p>It worked for us 😊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332161,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/02/2021 00:26:26",
          "content": "<p>Seems to work the same<br>\n<img src=\"https://user-images.githubusercontent.com/16153860/120405827-6c007200-c384-11eb-86b7-456b2e686e9e.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332178,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "06/02/2021 00:40:57",
          "content": "<p>Thank you for the info for both of you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332193,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "06/02/2021 00:55:55",
          "content": "<p>I was stalling to implement a location based post processing for a long time, just the thought about dealing with name files on the private dataset was enough to make me focus on something else haha.</p>\n<p>Now that i know that it works i kind of want to implement it just to see the impact</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332157,
      "author_name": "lplenka",
      "author_url": "",
      "post_date": "06/02/2021 00:24:45",
      "content": "<p>Congratulations 🎉 <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> we also did a similar post processing. Good job with \"Inference from global information\" 👋</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332165,
      "author_name": "whutddmm",
      "author_url": "",
      "post_date": "06/02/2021 00:28:21",
      "content": "<p>Congratulations! 🎉🎉🎉🎉<br>\nWe also tried  the same post-processing by site location, but I used 80km, 200km, and 500km. Maybe this is the reason why I have no effect</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332168,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "06/02/2021 00:29:10",
      "content": "<p>Congrats on the solo gold and the 4th place!</p>\n<p>Thank you for sharing the interesting approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332180,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "06/02/2021 00:43:11",
      "content": "<p>Congratulations!<br>\nCould you please show me the code for the NormalizeMelSpec() function?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332186,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/02/2021 00:51:48",
          "content": "<p>Some of the mono_to_color functions used in the previous Birdcall competition have been applied to (bs, 1, mel, time).</p>\n<pre><code>class NormalizeMelSpec(nn.Module):\n    def __init__(self, eps=1e-6):\n        super().__init__()\n        self.eps = eps\n\n    def forward(self, X):\n        mean = X.mean((1, 2), keepdim=True)\n        std = X.std((1, 2), keepdim=True)\n        Xstd = (X - mean) / (std + self.eps)\n        norm_min, norm_max = Xstd.min(-1)[0].min(-1)[0], Xstd.max(-1)[0].max(-1)[0]\n        fix_ind = (norm_max - norm_min) &gt; self.eps * torch.ones_like(\n            (norm_max - norm_min)\n        )\n        V = torch.zeros_like(Xstd)\n        if fix_ind.sum():\n            V_fix = Xstd[fix_ind]\n            norm_max_fix = norm_max[fix_ind, None, None]\n            norm_min_fix = norm_min[fix_ind, None, None]\n            V_fix = torch.max(\n                torch.min(V_fix, norm_max_fix),\n                norm_min_fix,\n            )\n            # print(V_fix.shape, norm_min_fix.shape, norm_max_fix.shape)\n            V_fix = (V_fix - norm_min_fix) / (norm_max_fix - norm_min_fix)\n            V[fix_ind] = V_fix\n        return V\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332211,
          "author_name": "shigemitsutomizawa",
          "author_url": "",
          "post_date": "06/02/2021 01:22:15",
          "content": "<p>Thank you. I'll try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332181,
      "author_name": "nitindatta",
      "author_url": "",
      "post_date": "06/02/2021 00:44:29",
      "content": "<p>Big Congratulations on your solo gold…<br>\nIf you dont mind can you please share how long did it take for you train the SED model for 50 epochs and the hardware you used??</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332184,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "06/02/2021 00:50:49",
          "content": "<p>That's actually a great question, i tried to train a SED model in colab pro and they banned me for abuse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332189,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/02/2021 00:52:51",
          "content": "<p>I used 1080ti x 3, it takes about 10~15 hours to train 50epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332194,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "06/02/2021 00:57:02",
          "content": "<p>That's not much actually! Congratz for the clean code!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332406,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "06/02/2021 05:41:02",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> on 4th place and thanks for sharing solution </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332689,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "06/02/2021 08:42:36",
      "content": "<p>Congratulations , i've one question, did you use spec augmentations ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1332711,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/02/2021 08:54:46",
          "content": "<p>No, it didn't work for my solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1332730,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "06/02/2021 09:08:23",
      "content": "<p>Congrats to the score and great solution! 62 models, that’s an ensemble to remember!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1332747,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/02/2021 09:18:17",
      "content": "<p>Congrats on the result.  I'll reread your writeup as I am now convinced that training on 5 seconds clips as I did is not the way to go.  You made SED models work quite well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1334449,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "06/03/2021 14:01:56",
      "content": "<p>I have released the experimental code and the inference notebook for the best submission.</p>\n<p>Code: <a href=\"https://github.com/tattaka/birdclef-2021\" target=\"_blank\">https://github.com/tattaka/birdclef-2021</a><br>\nInference notebook: <a href=\"https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\" target=\"_blank\">https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465</a><br>\nAggregation of the number of birds for post-processing: <a href=\"https://www.kaggle.com/tattaka/make-month-and-site-mask\" target=\"_blank\">https://www.kaggle.com/tattaka/make-month-and-site-mask</a></p>\n<p>In cleaning the code, I've removed duplicate experiments and non-reproducible ones that didn't specify a seed, but you should get similar scores.<br>\nPlease forgive me if my inference notebook is not organized due to laziness and time-saving.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1346879,
      "author_name": "mariotsaberlin",
      "author_url": "",
      "post_date": "06/12/2021 18:11:17",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> on 4th place and thanks for sharing your code and notebooks! <br>\nI was wondering about your thresholds. Do you have an explanation why your optimal threshold (around 0.05) is so much lower, compared to other solutions (usually around 0.3 even for other SED approaches)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1350497,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/15/2021 13:35:06",
          "content": "<p>With standard 5s prediction, the optimal threshold was close to 0.2.<br>\nMaybe, the use of global and local prediction can be used to stabilize and increase the sensitivity of your inferences.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1348410,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "06/14/2021 03:26:33",
      "content": "<p>Thank you for sharing!</p>\n<p>I have a question about \"Inference with global information\".<br>\nI checked your code. And my recognition is </p>\n<ul>\n<li>get long clip feat (eg. 10sec)</li>\n<li>get feat[0sec:5sec] and feat[5sec:10sec]</li>\n<li>get clipwise_prediction(feat[0:5]) and clipwise_prediction(feat[5:10])</li>\n</ul>\n<p>The effect of your idea is that we can get information about <strong>the seams</strong> of the clip.<br>\nIs this this recognition correct?</p>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/6258610a-d947-fb75-acb3-981b9db206c0.png\" alt=\"image.png\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1350519,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "06/15/2021 14:01:10",
          "content": "<p>Instead of getting two predictions from 10s to 5s, I get predictions from the central 5s and the whole 10s. Therefore, the number of inferences is the same as the normal 5s prediction.<br>\nThe disadvantage is that it takes longer to create a logmelspec than with 5s.   <br>\nAlso, inference near the edge of the clip repeats the clip to make up for the missing.<br>\nSo the global inference at the edge of the clip may be incorrect: (<br>\n<img src=\"https://user-images.githubusercontent.com/16153860/122066104-68adc180-ce2d-11eb-98b1-8869ba7cd4bd.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1351033,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "06/16/2021 03:03:26",
          "content": "<p>Thanks for the reply.</p>\n<p>I had misunderstood.<br>\nBut your diagram helped me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1332153": "First of all, I would like to thank the organizers of the competition and the Kaggle team for the competition hosting. I would also like to thank all the participants who have shared their knowledge so generously, including the previous Birdcall competitions. Two weeks before the end of this competition, I didn't expect it to be so hard competition.\n\n## Keys of my solution\n* Ensemble\n  * Max 62model (Best private: 47model)\n<!-- * global情報を用いたinference -->\n* **Inference with global information for SED model**\n* Post-processing\n\n## Training\n### Preprocess\n<!-- logmelspectrogramはTorchAudioを用いて計算され、それぞれの平均と分散を用いて正規化されました。 -->\nThe logmelspectrograms were calculated using TorchAudio and normalised using the means and variances per a image.\n``` python\nlogmelspec_extractor = nn.Sequential(\n            MelSpectrogram(\n                32000,\n                n_mels=128,\n                f_min=20,\n                n_fft=2048,\n                hop_length=512,\n                normalized=True,\n            ),\n            AmplitudeToDB(top_db=80.0),\n            NormalizeMelSpec(),\n        )\n```  \n<!-- waveformでのAugmentationとして、pink noiseやgaussian noiseを加えました。 -->\nFor Augmentation in waveform, gaussian and uniform, pink noise are added random.  \n<!-- logmelspec上のaugmentationとして、Mixupを用いました。パラメータとして0.2~0.5の確率、alpha=0.8を用いました。 -->\nAs an augmentation on logmelspec, I used Mixup. As a parameter,  I used a probability of 0.2~0.5 and alpha=0.8.\n### Modeling\n<!-- araiさんのkernelと同じくSED modelを用いました。 -->\nI used the SED model like [@hidehisaarai1213 's kernel](https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter).\n<!-- 様々なbackboneを用い、それぞれbatchsizeが36になるように使用する秒数が調整されました(最大で30s, 最小で10s) -->\nUsing various backbones, the seconds(max 30s, min 10s) used training has been adjusted so that each has a batchsize of 36.   \n### Training\n<!-- @hidehisaarai1213のmodelとはloss関数の計算が少し異なり`clipwise_pred`を直接最適化しています。こうすることでbackward時のNaNの発生を抑制することができました。 -->\n`clipwise_pred` is optimized directly. By doing this, I was able to suppress the generation of NaN on backward.\n``` python\nloss = bce_with_logits(torch.logit(clipwise_pred), target) + 0.5 * bce_with_logits((framewise_logit.max(1)[0], target)\n```\n<!-- loss関数はnn.BCEWithLogitsLossを用いました。 -->\n<!-- The loss function used is nn.BCEWithLogitsLoss. -->\n\nOther points are that\n* use a secondary label\n  * soft labels such as 0.5 did not work\n* trained 40~50 epochs \n  \n### Pseudo labeling\n<!-- 私は以下の手順を組み合わせていくつかのパターンを試してみました。 -->\nI have tried some patterns using a combination of the following steps.\n<!-- * 単純にaudio file単位のrelabeling\n  * clipwise_predを用い、閾値は0.2を用いました -->\n* Simple per-audio file relabeling\n  * Use clipwise_pred and threshold = 0.2\n<!-- * 前回のコンペのaraiさんのsolutionにインスピレーションを受け、audio file全体を入力としたframewise_predとtime_attを保存し、学習に用いる期間分を切り取りclipwise_predを計算しました。 -->\n* Inspired by [@hidehisaarai1213 's solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183204) in the last competition, I saved `framewise_pred` and `time_att` with the whole audio file as input, and computed `clipwise_pred` by cropping the period used for training.\n<!-- * secondary_labelがないaudio fileだけ2番目の方法でpseudo labelを作成する -->\n* Create a pseudo label using the second method only for audio files without a secondary_label\n<!-- * 閾値がある一定より下(0.05)だった場合、0.1の確率でprimary labelとsecondary labelを削除しました。 -->\n* If the pseudo label was below a certain value (0.05), the probability of 0.1 was used to set the primary and secondary labels to 0.\n\n<!-- どのパターンも個々のモデルの性能は0.01ほど向上し、train_soundscapeでのscoreはensembleに貢献しませんでしたが、public LB上では多少機能しました。 -->\nAll patterns improved the performance of the single models by about 0.01, and although the score on train_soundscape did not contribute to ensemble, it did work somewhat on the public LB.\n\n### 30s finetuning\n<!-- 大きいモデルではbatchsizeを確保するために10sなどの短いsegmentで学習されています。そのため、ラベルがノイジーになります。これを解決するためにAttention moduleのみを30sのsegmentを用いて10epoch追加で学習させました。train_soundscapeでのscoreでわずかな改善があり、ensembleの中に含まれています。 -->\nLarger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs. There is a slight improvement in score with train_soundscape, and it is included in ensemble.\n\n### Checkpoint selection\n<!-- それぞれvalidationでもtrainingと同じ秒数で切り取ったsegmentを用いout-of-foldでのlossが最も小さいものを選びました。 -->\nThe checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.\n\n## Inference with global information\n<!-- 自分のsolutionのなかで最も効果的で気に入っている部分です。   -->\nThis is my favorite and most effective part of my solution.  \n<!-- はじめに、このcompetitionに参加する前から持っていたアイディアは、SEDモデルが正確に\"Sound Event Detection\"を実行できるのであれば、長い期間に対して行われた予測の中で必要な部分だけを切り取ることで短い期間の予測を改善することができるのではないかということでした。 -->\nTo begin with, the idea that I had before joining this competition was that if the SED model could accurately perform \"Sound Event Detection\", then it would be possible to improve predictions for the shorter time segment by cropping from the longer time segment the necessary parts of predictions.\n\n\n<!-- このようにすることで、より長期間の特徴を推測に用いることができます。 -->\nIn this way, longer term features can be used for inference.\n```python\nframewise_pred_5s = self.fix_scale(feat[:, :, start:end])\natt_5s = torch.softmax(time_att[:, :, start:end], dim=-1)\nclipwise_pred_5s = torch.sum(torch.sigmoid(framewise_pred_5s) * att_5s, dim=-1,)\n```\n<!-- このアイディアを実行することで0.03ほどtrain_soundscapeでのscoreとpublic LB共に改善することができます。実際には長い期間のsegmentとして30sを用い、予測したい5sの区間が中心にくるように実装しました。このアイディアはpseudo labelを作る2つめのパターンでも用いられています。   -->\nBy implementing this idea, I can improve both score and public LB in train_soundscape by about 0.03. In practice, I used 30s as the segment for the longer period and implemented it so that the 5s interval I want to predict is in the center. This idea is also used to create a pseudo label in the second pattern.  \n<!-- 次に、前回の[Birdcallでの7th解法](https://www.kaggle.com/c/birdsong-recognition/discussion/183571)での観察ではaudio file単位で高い閾値を使い出現することを確認できた鳥は、短いsegmentにおいても高い確率で出現することが期待でき、閾値を下げて感度を上げることができます。 -->\nSecondly, in [7th place solution in previous Birdcall competition](https://www.kaggle.com/c/birdsong-recognition/discussion/183571),  the birds that were observed to appear using a high threshold for each audio file can be expected to appear in short segments with a high probability, and the sensitivity can be increased by lowering the threshold.  \n<!-- この観察に触発され、`clipwise_pred_30s`では高い閾値(0.05)を用いて考えられる鳥のリストを作成し、`clipwise_pred_5s`では低い閾値(0.025)を用いAND演算を行いました。 -->\nInspired by this observation, I used a high threshold (0.05) in `clipwise_pred_30s` to generate a list of possible birds and a low threshold (0.025) in `clipwise_pred_5s`,  to perform AND operations.\n```python\n((clipwise_pred_30s > high_threshold) + (clipwise_pred_5s > low_threshold)) >= 2\n```\n<!-- これらの閾値の見当をつけるために`scipy.optimize.dual_annealing`を用い、train_soundscapeに対して最適化しました。これはかなり危険なように見えますが、後に述べる単純なpublic LB probeによりある程度妥当だと自分は考えました。 -->\nTo get an idea of these thresholds, I used `scipy.optimize.dual_annealing` to optimize for train_soundscape. This seems rather risky, but I thought it was somewhat reasonable due to the simple public LB probing described below.\n<!-- この二重の閾値処理によってpublic LBとtrain_soundscapeのscoreは0.03改善されました。   -->\nThis double thresholding improved the score of public LB and train_soundscape by 0.03.  \n<!-- アンサンブルでは単純な平均を取りました。初期にhard voteも検討しましたが、あまりスコアは変わらないため、単純な手法を選びました。   -->\nIn the ensemble I took a simple average. Initially, hard voting was considered, but the score did not change much, so the simple method was chosen.  \n<!-- ## 場所と日付を用いたpostprocessing -->\n## Post-processing with location and date\n<!-- それぞれのsite周辺450メートル内で観測された鳥のリストと、それぞれの月ごとに出現する鳥のリストを作成し、出現しないものをsubmissionから削除しました。 -->\nA list of birds observed within 450 meters around each site and a list of birds appearing in each month was made, and those not appearing were removed from the submission.\n<!-- ただし、\"COR\"と\"COL\"の2つのsiteは比較的近くのため、どちらも0の場合のみ削除されました。 -->\nHere, because the two sites \"COR\" and \"COL\" are relatively close, the birds have only been removed if both are zero.\n<!-- また、site周辺で観察された鳥の中でいくつかの希少なクラスに関して閾値を4倍にしました。   -->\nI have also quadrupled the thresholds for some rare classes of birds observed in the vicinity of the sites.\n<!-- これらのpostprocessingにより、train_soundscape・public LB共に0.01以下の一貫した改善が見られました。 -->\nThis post-processing resulted in a consistent improvement of less than 0.01 for both train_soundscape and public LB.\n\n## Simple public LB probing\n<!-- ある程度train_soundscapeとpublic LBは相関していたため、public LBの分布が極端にtrain_soundscapeと似ているのではないかという危惧を感じました。 -->\nSince train_soundscape and public LB were correlated to some extent, I was concerned that the distribution of public LB might be extremely similar to train_soundscape.\n<!-- それを検証するために、適当なsubmissionに対して、train_soundscapeに含まれていないsiteと鳥の予測をnobirdに変えてsubmitしました。 -->\nTo verify this, I changed the site and bird predictions not included in the train_soundscape to nobird.\n<!-- ここで変化がなければtrain_soundscapeととpublic LBは同じ分布であると考えられますが、かなりscoreが低下しました。(0.71->0.58) -->\nIf there is no change in the score, it is assumed that train_soundscape and public LB have the same distribution, but the score has decreased considerably. (bird: 0.71->0.64, site: 0.71->0.58)\n<!-- よってpublic LBにはtrain_soundscapeに含まれないsite・鳥が含まれることがわかります。 -->\nSo I can see that public LB contains sites and birds that are not included in train_soundscape.\n<!-- ここから、強引な考えですが、publicとprivate はランダムに分割されていると予測しました。(これは厳密には不確実ではありますが、他に手のうちようがありませんでした) -->\nFrom this, I forcefully predicted that public and private would be randomly split. (This is strictly uncertain, but there was nothing else I could do.)\n\n## Experimental code and the inference notebook for best submission\nCode: https://github.com/tattaka/birdclef-2021\nInference notebook: https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\nAggregation of the number of birds for post-processing: https://www.kaggle.com/tattaka/make-month-and-site-mask",
    "1332155": "Congrats for the solo gold and thank you for sharing your unique approach. Does the post-processing also work for private lb?",
    "1332157": "Congratulations 🎉 @tattaka we also did a similar post processing. Good job with \"Inference from global information\" 👋",
    "1332159": "It worked for us 😊",
    "1332161": "Seems to work the same\n![](https://user-images.githubusercontent.com/16153860/120405827-6c007200-c384-11eb-86b7-456b2e686e9e.png)",
    "1332165": "Congratulations! 🎉🎉🎉🎉\nWe also tried  the same post-processing by site location, but I used 80km, 200km, and 500km. Maybe this is the reason why I have no effect",
    "1332168": "Congrats on the solo gold and the 4th place!\n\nThank you for sharing the interesting approach.",
    "1332178": "Thank you for the info for both of you.",
    "1332180": "Congratulations!\nCould you please show me the code for the NormalizeMelSpec() function?",
    "1332181": "Big Congratulations on your solo gold...\nIf you dont mind can you please share how long did it take for you train the SED model for 50 epochs and the hardware you used??",
    "1332184": "That's actually a great question, i tried to train a SED model in colab pro and they banned me for abuse",
    "1332186": "Some of the mono_to_color functions used in the previous Birdcall competition have been applied to (bs, 1, mel, time).\n``` python\nclass NormalizeMelSpec(nn.Module):\n    def __init__(self, eps=1e-6):\n        super().__init__()\n        self.eps = eps\n\n    def forward(self, X):\n        mean = X.mean((1, 2), keepdim=True)\n        std = X.std((1, 2), keepdim=True)\n        Xstd = (X - mean) / (std + self.eps)\n        norm_min, norm_max = Xstd.min(-1)[0].min(-1)[0], Xstd.max(-1)[0].max(-1)[0]\n        fix_ind = (norm_max - norm_min) > self.eps * torch.ones_like(\n            (norm_max - norm_min)\n        )\n        V = torch.zeros_like(Xstd)\n        if fix_ind.sum():\n            V_fix = Xstd[fix_ind]\n            norm_max_fix = norm_max[fix_ind, None, None]\n            norm_min_fix = norm_min[fix_ind, None, None]\n            V_fix = torch.max(\n                torch.min(V_fix, norm_max_fix),\n                norm_min_fix,\n            )\n            # print(V_fix.shape, norm_min_fix.shape, norm_max_fix.shape)\n            V_fix = (V_fix - norm_min_fix) / (norm_max_fix - norm_min_fix)\n            V[fix_ind] = V_fix\n        return V\n\n```",
    "1332189": "I used 1080ti x 3, it takes about 10~15 hours to train 50epoch.",
    "1332193": "I was stalling to implement a location based post processing for a long time, just the thought about dealing with name files on the private dataset was enough to make me focus on something else haha.\n\nNow that i know that it works i kind of want to implement it just to see the impact",
    "1332194": "That's not much actually! Congratz for the clean code!",
    "1332211": "Thank you. I'll try.",
    "1332406": "Congrats @tattaka on 4th place and thanks for sharing solution",
    "1332689": "Congratulations , i've one question, did you use spec augmentations ?",
    "1332711": "No, it didn't work for my solution.",
    "1332730": "Congrats to the score and great solution! 62 models, that’s an ensemble to remember!",
    "1332747": "Congrats on the result.  I'll reread your writeup as I am now convinced that training on 5 seconds clips as I did is not the way to go.  You made SED models work quite well.",
    "1334449": "I have released the experimental code and the inference notebook for the best submission.\n\nCode: https://github.com/tattaka/birdclef-2021\nInference notebook: https://www.kaggle.com/tattaka/birdclef2021-submissions-pp-ave?scriptVersionId=64016465\nAggregation of the number of birds for post-processing: https://www.kaggle.com/tattaka/make-month-and-site-mask\n\nIn cleaning the code, I've removed duplicate experiments and non-reproducible ones that didn't specify a seed, but you should get similar scores.\nPlease forgive me if my inference notebook is not organized due to laziness and time-saving.",
    "1346879": "Congratulations @tattaka on 4th place and thanks for sharing your code and notebooks! \nI was wondering about your thresholds. Do you have an explanation why your optimal threshold (around 0.05) is so much lower, compared to other solutions (usually around 0.3 even for other SED approaches)?",
    "1348410": "Thank you for sharing!\n\nI have a question about \"Inference with global information\".\nI checked your code. And my recognition is \n+ get long clip feat (eg. 10sec)\n+ get feat[0sec:5sec] and feat[5sec:10sec]\n+ get clipwise_prediction(feat[0:5]) and clipwise_prediction(feat[5:10])\n\nThe effect of your idea is that we can get information about **the seams** of the clip.\nIs this this recognition correct?\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/6258610a-d947-fb75-acb3-981b9db206c0.png)",
    "1350497": "With standard 5s prediction, the optimal threshold was close to 0.2.\nMaybe, the use of global and local prediction can be used to stabilize and increase the sensitivity of your inferences.",
    "1350519": "Instead of getting two predictions from 10s to 5s, I get predictions from the central 5s and the whole 10s. Therefore, the number of inferences is the same as the normal 5s prediction.\nThe disadvantage is that it takes longer to create a logmelspec than with 5s.   \nAlso, inference near the edge of the clip repeats the clip to make up for the missing.\nSo the global inference at the edge of the clip may be incorrect: (\n![](https://user-images.githubusercontent.com/16153860/122066104-68adc180-ce2d-11eb-98b1-8869ba7cd4bd.png)",
    "1351033": "Thanks for the reply.\n\nI had misunderstood.\nBut your diagram helped me."
  },
  "source": "meta"
}