{
  "id": 183204,
  "title": "6th place solution and some thoughts",
  "url": "/competitions/birdsong-recognition/discussion/183204",
  "author_name": "Hidehisa Arai",
  "post_date": "2020-09-16T00:27:55.695000",
  "votes": 162,
  "comment_count": 64,
  "views": 0,
  "content": "<p>First of all, I would like to sincerely thank <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>, and the members of kaggle team for hosting this competition. I had a lot of fun tackling on some of the challenging problems of machine learning thinking of the generative process of the data. Also many thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for actively sharing a lot of deep insights. It helped me a lot to come up with some good ideas and also made me convinced that I was in good direction.</p>\n<p>Following the recent two competitions: <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment\" target=\"_blank\">PANDA</a> challenge and <a href=\"https://www.kaggle.com/c/global-wheat-detection\" target=\"_blank\">GWD</a> challenge, this competition was also about <strong>domain shift</strong> and <strong>noisy labels</strong>.<br>\nCombination of these two challenging topics made this competition extremely difficult and we were troubled a lot how to make stable validation scheme. To be honest, contrary to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>'s <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/181499#1010570\" target=\"_blank\">expectation</a>, I couldn't find any good local validation scheme as the labels of training dataset contains a lot of noise. <code>Trust LB</code> was also not a very good policy since public LB was only 27% and we didn't know how the test set was devided. Instead, I took the policy of <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230\" target=\"_blank\">PANDA's competitors</a> - <strong>Ignore CV, care about public LB, and trust methodology</strong>.</p>\n<p>Although this competition was difficult as a data science competition, it was actually pretty close to a realistic setting and was full of problems that we often face in applying data science to real-world problem. Especially the combination of domain shift and noisy labels often happens (I think) when we are to use data from User Generated Contents(UGC) web service like Xeno Canto, YouTube, Twitter for training machine learning algorithms. Therefore, I think my solution is useful not only for this competition but also for those data science tasks related with UGC data, as it's basically focused on dealing with noisy labels and domain shift.</p>\n<h2>Solution in three lines</h2>\n<ul>\n<li>3 stages of training to gradually remove noise in labels</li>\n<li>SED style training and inference as I introduced <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">here</a></li>\n<li>Ensemble of 11 EMA models trained with the whole dataset / whole <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">extended dataset</a> to stabilize the result</li>\n</ul>\n<p>Code is available here: <a href=\"https://github.com/koukyo1994/kaggle-birdcall-6th-place\" target=\"_blank\">koukyo1994/kaggle-birdcall-6th-place</a>.<br>\nNote: I've re-implemented the code from original repository since it's quite messy. However, I haven't checked whether whole pipeline works; if you find something wrong with the code, please let me know :)</p>\n<h2>Motivation</h2>\n<p>It's quite obvious that there is a huge gap between training dataset and test dataset: it was described by the host and can also be seen through submission. However, <em>how</em> they are different is not obvious. Therefore, the first thing I did was to understand by observation what kind of domain shifts are seen between training and test datasets.</p>\n<p>Domain shift is an umbrella term and there are several problem classes of domain shift. The most famous one is <em>covariate shift</em>, where \\( P_{train}(X)\\neq P_{test}(X) \\) and \\( P_{train}(Y|X)=P_{test}(Y|X) \\). This one is well-studied and several algorithms are proposed to deal with the situation, but is not the type of domain shift in this competition. Another problem class is <em>prior probability shift</em> or <em>target shift</em>, where \\( P_{train}(Y)\\neq P_{test}(Y) \\) and \\( P_{train}(X|Y)=P_{test}(X|Y) \\), also not the type in this competition because \\( P_{train}(X|Y) \\) is not the same as \\( P_{test}(X|Y) \\). In fact, in this competition, multiple distribution shifts are present - shift in input space, shift in prior probability of labels, and shift in the function which connects \\( X \\) and \\( Y \\).</p>\n<p>How should we tackle a problem with various distribution shifts? The answer is simple - <em>divide the difficulty</em>. As I wrote above, shifts were three folds:</p>\n<p>1) shift in input space. For example, SNR difference or difference in sound collection environment (device/sampling rate/temperature/weather/…) between train and test. Occurence of non-target sound events is also a part of this shift.<br>\n2) shift in prior probability of labels. Distribution difference of species or distribution difference of calltypes, or else.<br>\n3) shift in the function which connects \\( X \\) and \\( Y \\). This has a very strong relation with label noise. Thinking of how the labels were created in train dataset and in test dataset, one could come up with the fact that Label Function(LF) of train dataset is completely different from that of test dataset. The former is annotations of the uploader (and can have large variation), whereas the latter is probably those of dedicated annotator(s) (and possibly have smaller variation).</p>\n<p>I decided to address these one by one and applied the following techniques.</p>\n<ul>\n<li>For 1), providing all the possible variation for train dataset may help. This is done by data augmentation.</li>\n<li>For 2), I just couldn't come up with smart ideas. I used ensemble of multiple models trained with datasets with different distributions of the labels to address this, but I think that is suboptimal.</li>\n<li>For 3), correcting the labels of train dataset to make train LF closer to test LF can help.</li>\n</ul>\n<p>On the other hand, label noise is also a term that contains multiple problem classes. First, in this competition, labels of train dataset are provided as <em>weak labels</em>. As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> pointed out <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/174774#972122\" target=\"_blank\">here</a>, weak label can be treated as noisy label if we change the point of view. Also there are some missing labels, which I'll explain later.</p>\n<p>I also decided to address these one by one.</p>\n<ul>\n<li>For weak label as noisy label, I first train a model with long chunk and then use the prediction as corrected label. I at first tried to create strong labels but couldn't make the first stage model good enough, therefore I used the prediction to correct weak labels we have in train dataset.</li>\n<li>For missing label, I also used the prediction of a model to find those.</li>\n</ul>\n<h2>First stage - build a model useful enough for addressing missing labels</h2>\n<p>In this stage, I used PANNs model. I used some basic augmentations (<code>NoiseInjection</code>, <code>PitchShift</code>, <code>RandomVolume</code>) and used <code>secondary_labels</code>. The keys in this stage were two folds:</p>\n<ul>\n<li>train with long chunk(30s) so that it would include call events of the species in <code>primaly_label</code> and <code>secondary_labels</code></li>\n<li>use attention pooling and max pooling to get weak prediction from <code>framewise_output</code></li>\n</ul>\n<p>Here are the reason behind.</p>\n<h3>train with long chunk</h3>\n<p>Assume we have <code>primary_label</code> of <code>birdA</code> and <code>secondary_labels</code> of <code>birdB</code> and <code>birdC</code>. Melspectrogram of the corresponding audio clip is something like the figure below (sorry for my poor drawing). If we use small window size, it may not include any sound events or include some sound events but not enough for the given labels. To make the model learn correctly, we need to make each label correspond to call event(s) of each species. For this reason, I used long chunk. Maybe it is better to use longer chunk like 1 minutes or more but I compromised to use 30s chunk considering the time for computation.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1489975%2Fcd841cb51b7b1fa4de6a6b2b7880e2d6%2FIMG_BBFA97950653-1.jpeg?generation=1600216062555033&amp;alt=media\" alt=\"\"></p>\n<h3>Combination of attention pooling and max pooling to get weak prediction</h3>\n<p>This is something I shared <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167611\" target=\"_blank\">here</a>.<br>\nIn the comments in <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">my SED notebook</a>, some said there were lots of false positives. This comes from the following reason.</p>\n<p>Weak prediction from attention pooling is made by applying self-attention filter on <code>framewise_outputs</code>. Now, we only have weak labels, we calculate the loss with weak predictions and weak labels. The gradient will be distributed to self-attention layer and pointwise classifier, but for pointwise classifier, strong supervision comes when self-attention put high probability at that point. Therefore, pointwise classifier are more likely to produce high probability value and rely on the attention layer to reduce false positives. This is not good when we are interested in the output of pointwise classifier(<code>framewise_outputs</code>).</p>\n<p>On the contrary, max pooling suppresses high probability values come out from pointwise classifier but it also has a defect that it is weak to impulse noise. Therefore, combination of max pooling and attention pooling can make balanced prediction and we can expect that be good prediction. For this reason, I used both but not combining the weak prediction from each aggregation but use each output and calculate loss for each, and sum them up. Therefore, the loss function I used in this stage was like this</p>\n<pre><code>bce = BCELoss()\nloss_att = bce(weak_pred_with_attention, label)\nloss_max = bce(weak_pred_with_maxpooling, label)\nloss = 1.0 * loss_att + 0.5 * loss_max\n</code></pre>\n<h3>Summarize this stage</h3>\n<ul>\n<li>Single PANNs model</li>\n<li>BCE on <code>clipwise_output</code> and also on maxpooled output.</li>\n<li>Adam + CosineAnnealing, 55epochs training</li>\n<li>train with randomly cropped 30s chunk</li>\n<li>validate on randomly cropped 30s chunk</li>\n<li>Augmentations on raw waveform<ul>\n<li><code>NoiseInjection</code> (max noise amplitude 0.04)</li>\n<li><code>PitchShift</code> (max pitch level 3)</li>\n<li><code>RandomVolume</code> (max db level 4)</li></ul></li>\n</ul>\n<h3>Get oof prediction and use it to find missing labels</h3>\n<p>With training procedure above, the model would get around 0.575 - 0.578. In fact, this is the weight I used in the public notebook.<br>\nI trained 5folds and got oof prediction on the whole training set. Then I used this oof prediction to find missing labels.</p>\n<p>Missing labels are more likely to be found from samples that does not have <code>secondary_labels</code>. It is up to the uploader to fill in <code>secondary_labels</code> or <code>background</code>, so some uploaders may not feel like to fill in those. Therefore, I picked samples without <code>secondary_labels</code> and used oof prediction of those to get additional labels if the probability of species that are not in their <code>primary_label</code> is over 0.9.</p>\n<h2>Second stage - build a model with additional labels to get stronger labels</h2>\n<p>In this stage, I used SED model with ResNeSt encoder. The difference between first stage and second stage is not that big - only the model, the existence of found labels (the labels obtained from the oof prediction of the first stage), and the input. I started to use 3channel input. The first channel was normal log-melspectrogram and the second channel was PCEN. The third channel was also log-melspectrogram but instead of using <code>librosa.power_to_db(melspec)</code>, I used <code>librosa.power_to_db(melspec ** 1.5)</code>. The idea of using different input for each channel comes from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170959\" target=\"_blank\">this post</a>. The chunk size is also reduced to 20s because of the GPU memory size limitation.</p>\n<p>I also changed the attention pooling slightly given the <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments\" target=\"_blank\">advice of </a><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> to use <code>torch.tanh</code> instead of <code>torch.clamp</code>.</p>\n<p>The result of this stage got around 0.60. I also trained 5folds model and this time I got oof prediction of <code>framewise_outputs</code>.</p>\n<p>To summarize,</p>\n<ul>\n<li>SED model with ResNeSt50 encoder, attention pooling head of PANNs (<code>torch.clamp</code> -&gt; <code>torch.tanh</code>)</li>\n<li>BCE on <code>clipwise_output</code> and also on maxpooled output</li>\n<li>Adam + CosineAnnealing, 75epochs training</li>\n<li>Add additional <code>secondary_labels</code> found in stage1</li>\n<li>3channels input - [normal logmel, PCEN, <code>librosa.power_to_db(melspec ** 1.5)</code>]</li>\n<li>train with randomly cropped 20s chunk</li>\n<li>validate on randomly cropped 30s chunk</li>\n<li>Augmentations on raw waveform<ul>\n<li><code>NoiseInjection</code> (max noise amplitude 0.04)</li>\n<li><code>PitchShift</code> (max pitch level 3)</li>\n<li><code>RandomVolume</code> (max db level 4)</li></ul></li>\n</ul>\n<h2>Third stage - build a model with oof framewise_outputs</h2>\n<p>Despite the use of training with long chunk size and missing labels, label noise problem is far from solved. SNR level to decide whether a call event was present is different between training set and test set. In fact, that was also different between samples of training set because basically each annotator (uploader) had their own labeling criteria. For this reason, I used the oof prediction of <code>framewise_outputs</code> of the second stage to further correct the labels of the training dataset. </p>\n<p>In this stage, I also cropped 20s chunk randomly to get a batch. At this time I also cropped the corresponding part of the predicted <code>framewise_outputs</code> and apply threshold on them (threshold varies from 0.3 - 0.7). After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction. Still, this prediction can be noisy and may contain false positives, I apply <code>logical_and</code> between predicted labels and provided labels. In this way, I got corrected chunk level label and use that for training.</p>\n<p>Here's the pseudo-code of the process above</p>\n<pre><code>y_batch = y[start_index:end_index]\nsoft_label_batch = soft_label[start_index_for_label:end_index_for_label]  # (n_frames, 264)\nthresholded = (soft_label_batch &gt;= threshold).astype(int)\n\nweak_pred = thresholded.max(axis=0)  # (264,)\ncorrected_label = np.logical_and(label, weak_pred)  # (264,)\n</code></pre>\n<p>In this stage, I also tried EfficientNet-B0 encoder and FocalLoss. Combination of ResNeSt encoder and FocalLoss didn't work well, whereas EffNet-B0 and FocalLoss worked well on public LB.</p>\n<p>All the other settings were the same as that of second stage. After this, I got around 0.61x score on public LB.</p>\n<p>To summarize,</p>\n<ul>\n<li>SED model with ResNeSt50 encoder, attention pooling head of PANNs (<code>torch.clamp</code> -&gt; <code>torch.tanh</code>) / EfficientNet-B0 encoder</li>\n<li>BCE on <code>clipwise_output</code> and also on maxpooled output / FocalLoss on <code>clipwise_output</code> and also on maxpooled output for EfficientNet-B0</li>\n<li>Adam + CosineAnnealing, 75epochs training</li>\n<li>Add additional <code>secondary_labels</code> found in stage1</li>\n<li>Correct labels using the prediction of stage2 model.</li>\n<li>3channels input - [normal logmel, PCEN, <code>librosa.power_to_db(melspec ** 1.5)</code>]</li>\n<li>train with randomly cropped 20s chunk</li>\n<li>validate on randomly cropped 30s chunk</li>\n<li>Augmentations on raw waveform<ul>\n<li><code>NoiseInjection</code> (max noise amplitude 0.04)</li>\n<li><code>PitchShift</code> (max pitch level 3)</li>\n<li><code>RandomVolume</code> (max db level 4)</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<p>With the corrected chunk level label, I trained the model with the whole dataset and use EMA model (using the implementation <a href=\"https://pytorch.org/docs/stable/optim.html#stochastic-weight-averaging\" target=\"_blank\">here</a>) for inference. This is a technique also used in <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172458\" target=\"_blank\">GWD competition</a>. </p>\n<p>I prepared models with different threshold (threshold on <code>framewise_outputs</code> of oof prediction) to make the model robust. Also important was the use of extended dataset. It doesn't get better result on public LB but when I used that for ensemble, it bumped up the score.</p>\n<h2>Things that didn't work for me</h2>\n<ul>\n<li>Mixup</li>\n<li>Calltype classification (781 class)</li>\n<li>noisy student training</li>\n<li>Larger models of efficientnet (b1, b2, b3…)</li>\n<li>etc…</li>\n</ul>\n<h2>Things that worked but doesn't make sense to me</h2>\n<p>When I change the PERIOD used for inference, the result greatly changed. Basically the larger value PERIOD is, the better the result. I just couldn't figure out why.</p>\n<h2>Things that I wanted to try but couldn't</h2>\n<p>There are lots of things I couldn't do due to the time limitation </p>\n<ul>\n<li>Use a variety of models</li>\n<li>Further refinement of the labels of train dataset/train extended dataset</li>\n<li>Use of auxiliary predictor to predict longitude/latitude/elevation and use the prediction to correct the main classifier</li>\n<li>Post-processing to refine prediction using species correlation information (I couldn't get the API key for ebird.org therefore I couldn't collect correlation information)</li>\n<li>Mixing background noise</li>\n<li>etc…</li>\n</ul>\n<h2>The thing that helped me a lot during competition</h2>\n<p>I created a simple streamlit app to check the audio data. I mainly used this to check the effect of augmentations or to check the quality of SED models prediction. This helped me a lot to figure out the major problems of this competition</p>\n<p><a href=\"https://github.com/koukyo1994/streamlit-audio\" target=\"_blank\">https://github.com/koukyo1994/streamlit-audio</a></p>\n<p>Later I learned <a href=\"https://www.kaggle.com/fkubota\" target=\"_blank\">@fkubota</a> also made an app with similar functionality. I didn't use this but it seems it's better than mine.</p>\n<p><a href=\"https://github.com/fkubota/spectrogram-tree\" target=\"_blank\">https://github.com/fkubota/spectrogram-tree</a></p>",
  "messages": [
    {
      "id": 1012159,
      "postDate": "2020-09-16T00:27:55.697Z",
      "content": "<p>First of all, I would like to sincerely thank <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>, and the members of kaggle team for hosting this competition. I had a lot of fun tackling on some of the challenging problems of machine learning thinking of the generative process of the data. Also many thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for actively sharing a lot of deep insights. It helped me a lot to come up with some good ideas and also made me convinced that I was in good direction.</p>\n<p>Following the recent two competitions: <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment\" target=\"_blank\">PANDA</a> challenge and <a href=\"https://www.kaggle.com/c/global-wheat-detection\" target=\"_blank\">GWD</a> challenge, this competition was also about <strong>domain shift</strong> and <strong>noisy labels</strong>.<br>\nCombination of these two challenging topics made this competition extremely difficult and we were troubled a lot how to make stable validation scheme. To be honest, contrary to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>'s <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/181499#1010570\" target=\"_blank\">expectation</a>, I couldn't find any good local validation scheme as the labels of training dataset contains a lot of noise. <code>Trust LB</code> was also not a very good policy since public LB was only 27% and we didn't know how the test set was devided. Instead, I took the policy of <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230\" target=\"_blank\">PANDA's competitors</a> - <strong>Ignore CV, care about public LB, and trust methodology</strong>.</p>\n<p>Although this competition was difficult as a data science competition, it was actually pretty close to a realistic setting and was full of problems that we often face in applying data science to real-world problem. Especially the combination of domain shift and noisy labels often happens (I think) when we are to use data from User Generated Contents(UGC) web service like Xeno Canto, YouTube, Twitter for training machine learning algorithms. Therefore, I think my solution is useful not only for this competition but also for those data science tasks related with UGC data, as it's basically focused on dealing with noisy labels and domain shift.</p>\n<h2>Solution in three lines</h2>\n<ul>\n<li>3 stages of training to gradually remove noise in labels</li>\n<li>SED style training and inference as I introduced <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">here</a></li>\n<li>Ensemble of 11 EMA models trained with the whole dataset / whole <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">extended dataset</a> to stabilize the result</li>\n</ul>\n<p>Code is available here: <a href=\"https://github.com/koukyo1994/kaggle-birdcall-6th-place\" target=\"_blank\">koukyo1994/kaggle-birdcall-6th-place</a>.<br>\nNote: I've re-implemented the code from original repository since it's quite messy. However, I haven't checked whether whole pipeline works; if you find something wrong with the code, please let me know :)</p>\n<h2>Motivation</h2>\n<p>It's quite obvious that there is a huge gap between training dataset and test dataset: it was described by the host and can also be seen through submission. However, <em>how</em> they are different is not obvious. Therefore, the first thing I did was to understand by observation what kind of domain shifts are seen between training and test datasets.</p>\n<p>Domain shift is an umbrella term and there are several problem classes of domain shift. The most famous one is <em>covariate shift</em>, where \\( P_{train}(X)\\neq P_{test}(X) \\) and \\( P_{train}(Y|X)=P_{test}(Y|X) \\). This one is well-studied and several algorithms are proposed to deal with the situation, but is not the type of domain shift in this competition. Another problem class is <em>prior probability shift</em> or <em>target shift</em>, where \\( P_{train}(Y)\\neq P_{test}(Y) \\) and \\( P_{train}(X|Y)=P_{test}(X|Y) \\), also not the type in this competition because \\( P_{train}(X|Y) \\) is not the same as \\( P_{test}(X|Y) \\). In fact, in this competition, multiple distribution shifts are present - shift in input space, shift in prior probability of labels, and shift in the function which connects \\( X \\) and \\( Y \\).</p>\n<p>How should we tackle a problem with various distribution shifts? The answer is simple - <em>divide the difficulty</em>. As I wrote above, shifts were three folds:</p>\n<p>1) shift in input space. For example, SNR difference or difference in sound collection environment (device/sampling rate/temperature/weather/…) between train and test. Occurence of non-target sound events is also a part of this shift.<br>\n2) shift in prior probability of labels. Distribution difference of species or distribution difference of calltypes, or else.<br>\n3) shift in the function which connects \\( X \\) and \\( Y \\). This has a very strong relation with label noise. Thinking of how the labels were created in train dataset and in test dataset, one could come up with the fact that Label Function(LF) of train dataset is completely different from that of test dataset. The former is annotations of the uploader (and can have large variation), whereas the latter is probably those of dedicated annotator(s) (and possibly have smaller variation).</p>\n<p>I decided to address these one by one and applied the following techniques.</p>\n<ul>\n<li>For 1), providing all the possible variation for train dataset may help. This is done by data augmentation.</li>\n<li>For 2), I just couldn't come up with smart ideas. I used ensemble of multiple models trained with datasets with different distributions of the labels to address this, but I think that is suboptimal.</li>\n<li>For 3), correcting the labels of train dataset to make train LF closer to test LF can help.</li>\n</ul>\n<p>On the other hand, label noise is also a term that contains multiple problem classes. First, in this competition, labels of train dataset are provided as <em>weak labels</em>. As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> pointed out <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/174774#972122\" target=\"_blank\">here</a>, weak label can be treated as noisy label if we change the point of view. Also there are some missing labels, which I'll explain later.</p>\n<p>I also decided to address these one by one.</p>\n<ul>\n<li>For weak label as noisy label, I first train a model with long chunk and then use the prediction as corrected label. I at first tried to create strong labels but couldn't make the first stage model good enough, therefore I used the prediction to correct weak labels we have in train dataset.</li>\n<li>For missing label, I also used the prediction of a model to find those.</li>\n</ul>\n<h2>First stage - build a model useful enough for addressing missing labels</h2>\n<p>In this stage, I used PANNs model. I used some basic augmentations (<code>NoiseInjection</code>, <code>PitchShift</code>, <code>RandomVolume</code>) and used <code>secondary_labels</code>. The keys in this stage were two folds:</p>\n<ul>\n<li>train with long chunk(30s) so that it would include call events of the species in <code>primaly_label</code> and <code>secondary_labels</code></li>\n<li>use attention pooling and max pooling to get weak prediction from <code>framewise_output</code></li>\n</ul>\n<p>Here are the reason behind.</p>\n<h3>train with long chunk</h3>\n<p>Assume we have <code>primary_label</code> of <code>birdA</code> and <code>secondary_labels</code> of <code>birdB</code> and <code>birdC</code>. Melspectrogram of the corresponding audio clip is something like the figure below (sorry for my poor drawing). If we use small window size, it may not include any sound events or include some sound events but not enough for the given labels. To make the model learn correctly, we need to make each label correspond to call event(s) of each species. For this reason, I used long chunk. Maybe it is better to use longer chunk like 1 minutes or more but I compromised to use 30s chunk considering the time for computation.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1489975%2Fcd841cb51b7b1fa4de6a6b2b7880e2d6%2FIMG_BBFA97950653-1.jpeg?generation=1600216062555033&amp;alt=media\" alt=\"\"></p>\n<h3>Combination of attention pooling and max pooling to get weak prediction</h3>\n<p>This is something I shared <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167611\" target=\"_blank\">here</a>.<br>\nIn the comments in <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">my SED notebook</a>, some said there were lots of false positives. This comes from the following reason.</p>\n<p>Weak prediction from attention pooling is made by applying self-attention filter on <code>framewise_outputs</code>. Now, we only have weak labels, we calculate the loss with weak predictions and weak labels. The gradient will be distributed to self-attention layer and pointwise classifier, but for pointwise classifier, strong supervision comes when self-attention put high probability at that point. Therefore, pointwise classifier are more likely to produce high probability value and rely on the attention layer to reduce false positives. This is not good when we are interested in the output of pointwise classifier(<code>framewise_outputs</code>).</p>\n<p>On the contrary, max pooling suppresses high probability values come out from pointwise classifier but it also has a defect that it is weak to impulse noise. Therefore, combination of max pooling and attention pooling can make balanced prediction and we can expect that be good prediction. For this reason, I used both but not combining the weak prediction from each aggregation but use each output and calculate loss for each, and sum them up. Therefore, the loss function I used in this stage was like this</p>\n<pre><code>bce = BCELoss()\nloss_att = bce(weak_pred_with_attention, label)\nloss_max = bce(weak_pred_with_maxpooling, label)\nloss = 1.0 * loss_att + 0.5 * loss_max\n</code></pre>\n<h3>Summarize this stage</h3>\n<ul>\n<li>Single PANNs model</li>\n<li>BCE on <code>clipwise_output</code> and also on maxpooled output.</li>\n<li>Adam + CosineAnnealing, 55epochs training</li>\n<li>train with randomly cropped 30s chunk</li>\n<li>validate on randomly cropped 30s chunk</li>\n<li>Augmentations on raw waveform<ul>\n<li><code>NoiseInjection</code> (max noise amplitude 0.04)</li>\n<li><code>PitchShift</code> (max pitch level 3)</li>\n<li><code>RandomVolume</code> (max db level 4)</li></ul></li>\n</ul>\n<h3>Get oof prediction and use it to find missing labels</h3>\n<p>With training procedure above, the model would get around 0.575 - 0.578. In fact, this is the weight I used in the public notebook.<br>\nI trained 5folds and got oof prediction on the whole training set. Then I used this oof prediction to find missing labels.</p>\n<p>Missing labels are more likely to be found from samples that does not have <code>secondary_labels</code>. It is up to the uploader to fill in <code>secondary_labels</code> or <code>background</code>, so some uploaders may not feel like to fill in those. Therefore, I picked samples without <code>secondary_labels</code> and used oof prediction of those to get additional labels if the probability of species that are not in their <code>primary_label</code> is over 0.9.</p>\n<h2>Second stage - build a model with additional labels to get stronger labels</h2>\n<p>In this stage, I used SED model with ResNeSt encoder. The difference between first stage and second stage is not that big - only the model, the existence of found labels (the labels obtained from the oof prediction of the first stage), and the input. I started to use 3channel input. The first channel was normal log-melspectrogram and the second channel was PCEN. The third channel was also log-melspectrogram but instead of using <code>librosa.power_to_db(melspec)</code>, I used <code>librosa.power_to_db(melspec ** 1.5)</code>. The idea of using different input for each channel comes from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170959\" target=\"_blank\">this post</a>. The chunk size is also reduced to 20s because of the GPU memory size limitation.</p>\n<p>I also changed the attention pooling slightly given the <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments\" target=\"_blank\">advice of </a><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> to use <code>torch.tanh</code> instead of <code>torch.clamp</code>.</p>\n<p>The result of this stage got around 0.60. I also trained 5folds model and this time I got oof prediction of <code>framewise_outputs</code>.</p>\n<p>To summarize,</p>\n<ul>\n<li>SED model with ResNeSt50 encoder, attention pooling head of PANNs (<code>torch.clamp</code> -&gt; <code>torch.tanh</code>)</li>\n<li>BCE on <code>clipwise_output</code> and also on maxpooled output</li>\n<li>Adam + CosineAnnealing, 75epochs training</li>\n<li>Add additional <code>secondary_labels</code> found in stage1</li>\n<li>3channels input - [normal logmel, PCEN, <code>librosa.power_to_db(melspec ** 1.5)</code>]</li>\n<li>train with randomly cropped 20s chunk</li>\n<li>validate on randomly cropped 30s chunk</li>\n<li>Augmentations on raw waveform<ul>\n<li><code>NoiseInjection</code> (max noise amplitude 0.04)</li>\n<li><code>PitchShift</code> (max pitch level 3)</li>\n<li><code>RandomVolume</code> (max db level 4)</li></ul></li>\n</ul>\n<h2>Third stage - build a model with oof framewise_outputs</h2>\n<p>Despite the use of training with long chunk size and missing labels, label noise problem is far from solved. SNR level to decide whether a call event was present is different between training set and test set. In fact, that was also different between samples of training set because basically each annotator (uploader) had their own labeling criteria. For this reason, I used the oof prediction of <code>framewise_outputs</code> of the second stage to further correct the labels of the training dataset. </p>\n<p>In this stage, I also cropped 20s chunk randomly to get a batch. At this time I also cropped the corresponding part of the predicted <code>framewise_outputs</code> and apply threshold on them (threshold varies from 0.3 - 0.7). After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction. Still, this prediction can be noisy and may contain false positives, I apply <code>logical_and</code> between predicted labels and provided labels. In this way, I got corrected chunk level label and use that for training.</p>\n<p>Here's the pseudo-code of the process above</p>\n<pre><code>y_batch = y[start_index:end_index]\nsoft_label_batch = soft_label[start_index_for_label:end_index_for_label]  # (n_frames, 264)\nthresholded = (soft_label_batch &gt;= threshold).astype(int)\n\nweak_pred = thresholded.max(axis=0)  # (264,)\ncorrected_label = np.logical_and(label, weak_pred)  # (264,)\n</code></pre>\n<p>In this stage, I also tried EfficientNet-B0 encoder and FocalLoss. Combination of ResNeSt encoder and FocalLoss didn't work well, whereas EffNet-B0 and FocalLoss worked well on public LB.</p>\n<p>All the other settings were the same as that of second stage. After this, I got around 0.61x score on public LB.</p>\n<p>To summarize,</p>\n<ul>\n<li>SED model with ResNeSt50 encoder, attention pooling head of PANNs (<code>torch.clamp</code> -&gt; <code>torch.tanh</code>) / EfficientNet-B0 encoder</li>\n<li>BCE on <code>clipwise_output</code> and also on maxpooled output / FocalLoss on <code>clipwise_output</code> and also on maxpooled output for EfficientNet-B0</li>\n<li>Adam + CosineAnnealing, 75epochs training</li>\n<li>Add additional <code>secondary_labels</code> found in stage1</li>\n<li>Correct labels using the prediction of stage2 model.</li>\n<li>3channels input - [normal logmel, PCEN, <code>librosa.power_to_db(melspec ** 1.5)</code>]</li>\n<li>train with randomly cropped 20s chunk</li>\n<li>validate on randomly cropped 30s chunk</li>\n<li>Augmentations on raw waveform<ul>\n<li><code>NoiseInjection</code> (max noise amplitude 0.04)</li>\n<li><code>PitchShift</code> (max pitch level 3)</li>\n<li><code>RandomVolume</code> (max db level 4)</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<p>With the corrected chunk level label, I trained the model with the whole dataset and use EMA model (using the implementation <a href=\"https://pytorch.org/docs/stable/optim.html#stochastic-weight-averaging\" target=\"_blank\">here</a>) for inference. This is a technique also used in <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172458\" target=\"_blank\">GWD competition</a>. </p>\n<p>I prepared models with different threshold (threshold on <code>framewise_outputs</code> of oof prediction) to make the model robust. Also important was the use of extended dataset. It doesn't get better result on public LB but when I used that for ensemble, it bumped up the score.</p>\n<h2>Things that didn't work for me</h2>\n<ul>\n<li>Mixup</li>\n<li>Calltype classification (781 class)</li>\n<li>noisy student training</li>\n<li>Larger models of efficientnet (b1, b2, b3…)</li>\n<li>etc…</li>\n</ul>\n<h2>Things that worked but doesn't make sense to me</h2>\n<p>When I change the PERIOD used for inference, the result greatly changed. Basically the larger value PERIOD is, the better the result. I just couldn't figure out why.</p>\n<h2>Things that I wanted to try but couldn't</h2>\n<p>There are lots of things I couldn't do due to the time limitation </p>\n<ul>\n<li>Use a variety of models</li>\n<li>Further refinement of the labels of train dataset/train extended dataset</li>\n<li>Use of auxiliary predictor to predict longitude/latitude/elevation and use the prediction to correct the main classifier</li>\n<li>Post-processing to refine prediction using species correlation information (I couldn't get the API key for ebird.org therefore I couldn't collect correlation information)</li>\n<li>Mixing background noise</li>\n<li>etc…</li>\n</ul>\n<h2>The thing that helped me a lot during competition</h2>\n<p>I created a simple streamlit app to check the audio data. I mainly used this to check the effect of augmentations or to check the quality of SED models prediction. This helped me a lot to figure out the major problems of this competition</p>\n<p><a href=\"https://github.com/koukyo1994/streamlit-audio\" target=\"_blank\">https://github.com/koukyo1994/streamlit-audio</a></p>\n<p>Later I learned <a href=\"https://www.kaggle.com/fkubota\" target=\"_blank\">@fkubota</a> also made an app with similar functionality. I didn't use this but it seems it's better than mine.</p>\n<p><a href=\"https://github.com/fkubota/spectrogram-tree\" target=\"_blank\">https://github.com/fkubota/spectrogram-tree</a></p>",
      "rawMarkdown": "First of all, I would like to sincerely thank [@stefankahl](https://www.kaggle.com/stefankahl), [@tomdenton](https://www.kaggle.com/tomdenton), [@holgerklinck](https://www.kaggle.com/holgerklinck), and the members of kaggle team for hosting this competition. I had a lot of fun tackling on some of the challenging problems of machine learning thinking of the generative process of the data. Also many thanks to [@hengck23](https://www.kaggle.com/hengck23) for actively sharing a lot of deep insights. It helped me a lot to come up with some good ideas and also made me convinced that I was in good direction.\n\nFollowing the recent two competitions: [PANDA](https://www.kaggle.com/c/prostate-cancer-grade-assessment) challenge and [GWD](https://www.kaggle.com/c/global-wheat-detection) challenge, this competition was also about **domain shift** and **noisy labels**.\nCombination of these two challenging topics made this competition extremely difficult and we were troubled a lot how to make stable validation scheme. To be honest, contrary to [@cpmpml](https://www.kaggle.com/cpmpml)'s [expectation](https://www.kaggle.com/c/birdsong-recognition/discussion/181499#1010570), I couldn't find any good local validation scheme as the labels of training dataset contains a lot of noise. `Trust LB` was also not a very good policy since public LB was only 27% and we didn't know how the test set was devided. Instead, I took the policy of [PANDA's competitors](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230) - **Ignore CV, care about public LB, and trust methodology**.\n\nAlthough this competition was difficult as a data science competition, it was actually pretty close to a realistic setting and was full of problems that we often face in applying data science to real-world problem. Especially the combination of domain shift and noisy labels often happens (I think) when we are to use data from User Generated Contents(UGC) web service like Xeno Canto, YouTube, Twitter for training machine learning algorithms. Therefore, I think my solution is useful not only for this competition but also for those data science tasks related with UGC data, as it's basically focused on dealing with noisy labels and domain shift.\n\n## Solution in three lines\n\n* 3 stages of training to gradually remove noise in labels\n* SED style training and inference as I introduced [here](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)\n* Ensemble of 11 EMA models trained with the whole dataset / whole [extended dataset](https://www.kaggle.com/c/birdsong-recognition/discussion/159970) to stabilize the result\n\nCode is available here: [koukyo1994/kaggle-birdcall-6th-place](https://github.com/koukyo1994/kaggle-birdcall-6th-place).\nNote: I've re-implemented the code from original repository since it's quite messy. However, I haven't checked whether whole pipeline works; if you find something wrong with the code, please let me know :)\n\n## Motivation\n\nIt's quite obvious that there is a huge gap between training dataset and test dataset: it was described by the host and can also be seen through submission. However, *how* they are different is not obvious. Therefore, the first thing I did was to understand by observation what kind of domain shifts are seen between training and test datasets.\n\nDomain shift is an umbrella term and there are several problem classes of domain shift. The most famous one is *covariate shift*, where \\\\( P_{train}(X)\\neq P_{test}(X) \\\\) and \\\\( P_{train}(Y|X)=P_{test}(Y|X) \\\\). This one is well-studied and several algorithms are proposed to deal with the situation, but is not the type of domain shift in this competition. Another problem class is *prior probability shift* or *target shift*, where \\\\( P_{train}(Y)\\neq P_{test}(Y) \\\\) and \\\\( P_{train}(X|Y)=P_{test}(X|Y) \\\\), also not the type in this competition because \\\\( P_{train}(X|Y) \\\\) is not the same as \\\\( P_{test}(X|Y) \\\\). In fact, in this competition, multiple distribution shifts are present - shift in input space, shift in prior probability of labels, and shift in the function which connects \\\\( X \\\\) and \\\\( Y \\\\).\n\nHow should we tackle a problem with various distribution shifts? The answer is simple - *divide the difficulty*. As I wrote above, shifts were three folds:\n\n1) shift in input space. For example, SNR difference or difference in sound collection environment (device/sampling rate/temperature/weather/...) between train and test. Occurence of non-target sound events is also a part of this shift.\n2) shift in prior probability of labels. Distribution difference of species or distribution difference of calltypes, or else.\n3) shift in the function which connects \\\\( X \\\\) and \\\\( Y \\\\). This has a very strong relation with label noise. Thinking of how the labels were created in train dataset and in test dataset, one could come up with the fact that Label Function(LF) of train dataset is completely different from that of test dataset. The former is annotations of the uploader (and can have large variation), whereas the latter is probably those of dedicated annotator(s) (and possibly have smaller variation).\n\nI decided to address these one by one and applied the following techniques.\n\n* For 1), providing all the possible variation for train dataset may help. This is done by data augmentation.\n* For 2), I just couldn't come up with smart ideas. I used ensemble of multiple models trained with datasets with different distributions of the labels to address this, but I think that is suboptimal.\n* For 3), correcting the labels of train dataset to make train LF closer to test LF can help.\n\nOn the other hand, label noise is also a term that contains multiple problem classes. First, in this competition, labels of train dataset are provided as *weak labels*. As @hengck23 pointed out [here](https://www.kaggle.com/c/birdsong-recognition/discussion/174774#972122), weak label can be treated as noisy label if we change the point of view. Also there are some missing labels, which I'll explain later.\n\nI also decided to address these one by one.\n\n* For weak label as noisy label, I first train a model with long chunk and then use the prediction as corrected label. I at first tried to create strong labels but couldn't make the first stage model good enough, therefore I used the prediction to correct weak labels we have in train dataset.\n* For missing label, I also used the prediction of a model to find those.\n\n## First stage - build a model useful enough for addressing missing labels\n\nIn this stage, I used PANNs model. I used some basic augmentations (`NoiseInjection`, `PitchShift`, `RandomVolume`) and used `secondary_labels`. The keys in this stage were two folds:\n\n* train with long chunk(30s) so that it would include call events of the species in `primaly_label` and `secondary_labels`\n* use attention pooling and max pooling to get weak prediction from `framewise_output`\n\nHere are the reason behind.\n\n### train with long chunk\n\nAssume we have `primary_label` of `birdA` and `secondary_labels` of `birdB` and `birdC`. Melspectrogram of the corresponding audio clip is something like the figure below (sorry for my poor drawing). If we use small window size, it may not include any sound events or include some sound events but not enough for the given labels. To make the model learn correctly, we need to make each label correspond to call event(s) of each species. For this reason, I used long chunk. Maybe it is better to use longer chunk like 1 minutes or more but I compromised to use 30s chunk considering the time for computation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1489975%2Fcd841cb51b7b1fa4de6a6b2b7880e2d6%2FIMG_BBFA97950653-1.jpeg?generation=1600216062555033&alt=media)\n\n### Combination of attention pooling and max pooling to get weak prediction\n\nThis is something I shared [here](https://www.kaggle.com/c/birdsong-recognition/discussion/167611).\nIn the comments in [my SED notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection), some said there were lots of false positives. This comes from the following reason.\n\nWeak prediction from attention pooling is made by applying self-attention filter on `framewise_outputs`. Now, we only have weak labels, we calculate the loss with weak predictions and weak labels. The gradient will be distributed to self-attention layer and pointwise classifier, but for pointwise classifier, strong supervision comes when self-attention put high probability at that point. Therefore, pointwise classifier are more likely to produce high probability value and rely on the attention layer to reduce false positives. This is not good when we are interested in the output of pointwise classifier(`framewise_outputs`).\n\nOn the contrary, max pooling suppresses high probability values come out from pointwise classifier but it also has a defect that it is weak to impulse noise. Therefore, combination of max pooling and attention pooling can make balanced prediction and we can expect that be good prediction. For this reason, I used both but not combining the weak prediction from each aggregation but use each output and calculate loss for each, and sum them up. Therefore, the loss function I used in this stage was like this\n\n```python\nbce = BCELoss()\nloss_att = bce(weak_pred_with_attention, label)\nloss_max = bce(weak_pred_with_maxpooling, label)\nloss = 1.0 * loss_att + 0.5 * loss_max\n```\n\n### Summarize this stage\n\n* Single PANNs model\n* BCE on `clipwise_output` and also on maxpooled output.\n* Adam + CosineAnnealing, 55epochs training\n* train with randomly cropped 30s chunk\n* validate on randomly cropped 30s chunk\n* Augmentations on raw waveform\n  - `NoiseInjection` (max noise amplitude 0.04)\n  - `PitchShift` (max pitch level 3)\n  - `RandomVolume` (max db level 4)\n\n### Get oof prediction and use it to find missing labels\n\nWith training procedure above, the model would get around 0.575 - 0.578. In fact, this is the weight I used in the public notebook.\nI trained 5folds and got oof prediction on the whole training set. Then I used this oof prediction to find missing labels.\n\nMissing labels are more likely to be found from samples that does not have `secondary_labels`. It is up to the uploader to fill in `secondary_labels` or `background`, so some uploaders may not feel like to fill in those. Therefore, I picked samples without `secondary_labels` and used oof prediction of those to get additional labels if the probability of species that are not in their `primary_label` is over 0.9.\n\n## Second stage - build a model with additional labels to get stronger labels\n\nIn this stage, I used SED model with ResNeSt encoder. The difference between first stage and second stage is not that big - only the model, the existence of found labels (the labels obtained from the oof prediction of the first stage), and the input. I started to use 3channel input. The first channel was normal log-melspectrogram and the second channel was PCEN. The third channel was also log-melspectrogram but instead of using `librosa.power_to_db(melspec)`, I used `librosa.power_to_db(melspec ** 1.5)`. The idea of using different input for each channel comes from [this post](https://www.kaggle.com/c/birdsong-recognition/discussion/170959). The chunk size is also reduced to 20s because of the GPU memory size limitation.\n\nI also changed the attention pooling slightly given the [advice of @hengck23](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments) to use `torch.tanh` instead of `torch.clamp`.\n\nThe result of this stage got around 0.60. I also trained 5folds model and this time I got oof prediction of `framewise_outputs`.\n\nTo summarize,\n\n* SED model with ResNeSt50 encoder, attention pooling head of PANNs (`torch.clamp` -> `torch.tanh`)\n* BCE on `clipwise_output` and also on maxpooled output\n* Adam + CosineAnnealing, 75epochs training\n* Add additional `secondary_labels` found in stage1\n* 3channels input - \\[normal logmel, PCEN, `librosa.power_to_db(melspec ** 1.5)`\\]\n* train with randomly cropped 20s chunk\n* validate on randomly cropped 30s chunk\n* Augmentations on raw waveform\n  - `NoiseInjection` (max noise amplitude 0.04)\n  - `PitchShift` (max pitch level 3)\n  - `RandomVolume` (max db level 4)\n\n## Third stage - build a model with oof framewise_outputs\n\nDespite the use of training with long chunk size and missing labels, label noise problem is far from solved. SNR level to decide whether a call event was present is different between training set and test set. In fact, that was also different between samples of training set because basically each annotator (uploader) had their own labeling criteria. For this reason, I used the oof prediction of `framewise_outputs` of the second stage to further correct the labels of the training dataset. \n\nIn this stage, I also cropped 20s chunk randomly to get a batch. At this time I also cropped the corresponding part of the predicted `framewise_outputs` and apply threshold on them (threshold varies from 0.3 - 0.7). After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction. Still, this prediction can be noisy and may contain false positives, I apply `logical_and` between predicted labels and provided labels. In this way, I got corrected chunk level label and use that for training.\n\nHere's the pseudo-code of the process above\n\n```python\ny_batch = y[start_index:end_index]\nsoft_label_batch = soft_label[start_index_for_label:end_index_for_label]  # (n_frames, 264)\nthresholded = (soft_label_batch >= threshold).astype(int)\n\nweak_pred = thresholded.max(axis=0)  # (264,)\ncorrected_label = np.logical_and(label, weak_pred)  # (264,)\n```\n\nIn this stage, I also tried EfficientNet-B0 encoder and FocalLoss. Combination of ResNeSt encoder and FocalLoss didn't work well, whereas EffNet-B0 and FocalLoss worked well on public LB.\n\nAll the other settings were the same as that of second stage. After this, I got around 0.61x score on public LB.\n\nTo summarize,\n\n* SED model with ResNeSt50 encoder, attention pooling head of PANNs (`torch.clamp` -> `torch.tanh`) / EfficientNet-B0 encoder\n* BCE on `clipwise_output` and also on maxpooled output / FocalLoss on `clipwise_output` and also on maxpooled output for EfficientNet-B0\n* Adam + CosineAnnealing, 75epochs training\n* Add additional `secondary_labels` found in stage1\n* Correct labels using the prediction of stage2 model.\n* 3channels input - \\[normal logmel, PCEN, `librosa.power_to_db(melspec ** 1.5)`\\]\n* train with randomly cropped 20s chunk\n* validate on randomly cropped 30s chunk\n* Augmentations on raw waveform\n  - `NoiseInjection` (max noise amplitude 0.04)\n  - `PitchShift` (max pitch level 3)\n  - `RandomVolume` (max db level 4)\n\n## Ensemble\n\nWith the corrected chunk level label, I trained the model with the whole dataset and use EMA model (using the implementation [here](https://pytorch.org/docs/stable/optim.html#stochastic-weight-averaging)) for inference. This is a technique also used in [GWD competition](https://www.kaggle.com/c/global-wheat-detection/discussion/172458). \n\nI prepared models with different threshold (threshold on `framewise_outputs` of oof prediction) to make the model robust. Also important was the use of extended dataset. It doesn't get better result on public LB but when I used that for ensemble, it bumped up the score.\n\n## Things that didn't work for me\n\n* Mixup\n* Calltype classification (781 class)\n* noisy student training\n* Larger models of efficientnet (b1, b2, b3...)\n* etc...\n\n## Things that worked but doesn't make sense to me\n\nWhen I change the PERIOD used for inference, the result greatly changed. Basically the larger value PERIOD is, the better the result. I just couldn't figure out why.\n\n## Things that I wanted to try but couldn't\n\nThere are lots of things I couldn't do due to the time limitation \n\n* Use a variety of models\n* Further refinement of the labels of train dataset/train extended dataset\n* Use of auxiliary predictor to predict longitude/latitude/elevation and use the prediction to correct the main classifier\n* Post-processing to refine prediction using species correlation information (I couldn't get the API key for ebird.org therefore I couldn't collect correlation information)\n* Mixing background noise\n* etc...\n\n## The thing that helped me a lot during competition\n\nI created a simple streamlit app to check the audio data. I mainly used this to check the effect of augmentations or to check the quality of SED models prediction. This helped me a lot to figure out the major problems of this competition\n\nhttps://github.com/koukyo1994/streamlit-audio\n\nLater I learned @fkubota also made an app with similar functionality. I didn't use this but it seems it's better than mine.\n\nhttps://github.com/fkubota/spectrogram-tree",
      "votes": 162
    },
    {
      "id": 1013294,
      "postDate": "2020-09-16T16:10:07.970Z",
      "content": "<p>Thanks for the extensive write-up and all of your engagement in the discussion forums! As you said, it's a hard task, and I don't think nearly as many people would have gotten a good start on it without your help.</p>",
      "rawMarkdown": "Thanks for the extensive write-up and all of your engagement in the discussion forums! As you said, it's a hard task, and I don't think nearly as many people would have gotten a good start on it without your help.",
      "votes": 7
    },
    {
      "id": 1137457,
      "postDate": "2021-01-04T00:33:10.763Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , thanks for sharing! I am learning what you shared here for the new Rainforest competition. :)<br>\nBut I have a question about the 3 channel inputs <code>[normal logmel, PCEN, librosa.power_to_db(melspec ** 1.5)]</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F65df09a5f2d60e6e8e7889f1355bed18%2FScreenshot%20from%202021-01-03%2018-28-45.png?generation=1609720253136614&amp;alt=media\" alt=\"\"></p>\n<p><code>power_to_db</code> is taking a log. Raising to a power before taking the log is just scaling the output. Can it really help the model? </p>",
      "rawMarkdown": "Hi, @hidehisaarai1213 , thanks for sharing! I am learning what you shared here for the new Rainforest competition. :)\nBut I have a question about the 3 channel inputs ` [normal logmel, PCEN, librosa.power_to_db(melspec ** 1.5)]`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F65df09a5f2d60e6e8e7889f1355bed18%2FScreenshot%20from%202021-01-03%2018-28-45.png?generation=1609720253136614&alt=media)\n\n`power_to_db` is taking a log. Raising to a power before taking the log is just scaling the output. Can it really help the model? ",
      "votes": 3
    },
    {
      "id": 1013417,
      "postDate": "2020-09-16T17:26:54.877Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for all your selfless sharing. <br>\nYou deserve a Gold just for that !</p>",
      "rawMarkdown": "Thank you @hidehisaarai1213 for all your selfless sharing. \nYou deserve a Gold just for that !",
      "votes": 3
    },
    {
      "id": 1012182,
      "postDate": "2020-09-16T00:41:19.177Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> on the gold medal. Thanks so much for sharing a detailed solution and for all your contributions in this cmpetition.</p>",
      "rawMarkdown": "Congrats @hidehisaarai1213 on the gold medal. Thanks so much for sharing a detailed solution and for all your contributions in this cmpetition.",
      "votes": 4
    },
    {
      "id": 1012177,
      "postDate": "2020-09-16T00:37:40.920Z",
      "content": "<p>Thanks a lot for all you've done for the competition, and congratz on the nice finish !</p>\n<p>Although I didn't believe in the SED idea (we didn't try it), we're glad that you made it work. I'll read your solution after a good night of sleep, giving it the full attention it deserves !</p>",
      "rawMarkdown": "Thanks a lot for all you've done for the competition, and congratz on the nice finish !\n\nAlthough I didn't believe in the SED idea (we didn't try it), we're glad that you made it work. I'll read your solution after a good night of sleep, giving it the full attention it deserves !",
      "votes": 4,
      "replies": [
        {
          "id": 1012221,
          "postDate": "2020-09-16T01:20:43.117Z",
          "content": "<p>Congratz again <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> !</p>\n<blockquote>\n  <p>Although I didn't believe in the SED idea (we didn't try it)</p>\n</blockquote>\n<p>Haha, I somehow knew this during the competition 😄</p>",
          "rawMarkdown": "Congratz again @theoviel !\n\n> Although I didn't believe in the SED idea (we didn't try it)\n\nHaha, I somehow knew this during the competition 😄",
          "votes": 1
        }
      ]
    },
    {
      "id": 1012174,
      "postDate": "2020-09-16T00:36:48.643Z",
      "content": "<p>Congrats again. I need to read this again in detail as it looks more elaborate than the noisy student method I wanted to try.</p>\n<p>I want to congratulate you for sharing so much as well.  </p>",
      "rawMarkdown": "Congrats again. I need to read this again in detail as it looks more elaborate than the noisy student method I wanted to try.\n\nI want to congratulate you for sharing so much as well.  ",
      "votes": 4,
      "replies": [
        {
          "id": 1012219,
          "postDate": "2020-09-16T01:16:45.120Z",
          "content": "<p>Thank you!</p>\n<blockquote>\n  <p>the noisy student method</p>\n</blockquote>\n<p>I actually thought of using noisy student method but couldn't make it work well 😑</p>\n<blockquote>\n  <p>I want to congratulate you for sharing so much as well.</p>\n</blockquote>\n<p>Thanks! This word is also for you! Thanks for actively sharing lots of things on discussion!</p>",
          "rawMarkdown": "Thank you!\n\n> the noisy student method\n\nI actually thought of using noisy student method but couldn't make it work well 😑\n\n> I want to congratulate you for sharing so much as well.\n\nThanks! This word is also for you! Thanks for actively sharing lots of things on discussion!",
          "votes": 2
        },
        {
          "id": 1012239,
          "postDate": "2020-09-16T01:34:41.733Z",
          "content": "<blockquote>\n  <p>I actually thought of using noisy student method but couldn't make it work well</p>\n</blockquote>\n<p>Good to know, I have no regrets then :D</p>",
          "rawMarkdown": "> I actually thought of using noisy student method but couldn't make it work well\n\nGood to know, I have no regrets then :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 1015686,
      "postDate": "2020-09-18T11:04:17.820Z",
      "content": "<p>Congratz and thank you for the summary.</p>\n<p>I was hoping for you all the best since the start of your competiton due to your great notebooks and explanations!</p>",
      "rawMarkdown": "Congratz and thank you for the summary.\n\nI was hoping for you all the best since the start of your competiton due to your great notebooks and explanations!",
      "votes": 1
    },
    {
      "id": 1015685,
      "postDate": "2020-09-18T11:04:04.380Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 1015408,
      "postDate": "2020-09-18T07:04:25.307Z",
      "content": "<p>Great solution!</p>",
      "rawMarkdown": "Great solution!",
      "votes": 1
    },
    {
      "id": 1015366,
      "postDate": "2020-09-18T06:27:54.567Z",
      "content": "<p>Thanks for writing this article. It really helps new folks like me out there!</p>",
      "rawMarkdown": "Thanks for writing this article. It really helps new folks like me out there!",
      "votes": 1
    },
    {
      "id": 1014280,
      "postDate": "2020-09-17T10:07:56.173Z",
      "content": "<p>That's what's up :)</p>",
      "rawMarkdown": "That's what's up :)",
      "votes": 1
    },
    {
      "id": 1014226,
      "postDate": "2020-09-17T09:17:40.747Z",
      "content": "<p>It's awesome. Thanks for being helpful to all by sharing your works. A lot to learn!</p>",
      "rawMarkdown": "It's awesome. Thanks for being helpful to all by sharing your works. A lot to learn!",
      "votes": 1
    },
    {
      "id": 1013274,
      "postDate": "2020-09-16T15:46:04.070Z",
      "content": "<p>Excellent good job</p>",
      "rawMarkdown": "Excellent good job",
      "votes": 1
    },
    {
      "id": 1013120,
      "postDate": "2020-09-16T14:11:35.930Z",
      "content": "<p>Congrats, great work and thanks so much for sharing during comp and now. </p>",
      "rawMarkdown": "Congrats, great work and thanks so much for sharing during comp and now. ",
      "votes": 1
    },
    {
      "id": 1012777,
      "postDate": "2020-09-16T09:31:32.187Z",
      "content": "<p>Excellent.. and great sharing. exposed some new concepts which will help me a lot. Thanks</p>",
      "rawMarkdown": "Excellent.. and great sharing. exposed some new concepts which will help me a lot. Thanks",
      "votes": 1
    },
    {
      "id": 1012655,
      "postDate": "2020-09-16T07:58:45.023Z",
      "content": "<p>You da real MVP <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> ! !</p>",
      "rawMarkdown": "You da real MVP @hidehisaarai1213 ! !",
      "votes": 1
    },
    {
      "id": 1012592,
      "postDate": "2020-09-16T07:06:45.043Z",
      "content": "<p>Congrats and excellent work!</p>\n<blockquote>\n  <p>and use EMA model for inference.</p>\n</blockquote>\n<p>Sorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.</p>",
      "rawMarkdown": "Congrats and excellent work!\n\n> and use EMA model for inference.\n\nSorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.",
      "votes": 1,
      "replies": [
        {
          "id": 1013036,
          "postDate": "2020-09-16T13:22:18.037Z",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Sorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.</p>\n</blockquote>\n<p>It's Exponential Moving Average in weight space. OK, I'll add some explanation :)</p>",
          "rawMarkdown": "Thanks!\n\n> Sorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.\n\nIt's Exponential Moving Average in weight space. OK, I'll add some explanation :)",
          "votes": 1
        },
        {
          "id": 1013902,
          "postDate": "2020-09-17T03:49:28.030Z",
          "content": "<p>Thanks for replying.<br>\nI read your solution again and have some questions. If you can talk more details that would be very nice.</p>\n<blockquote>\n  <p>After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction.</p>\n</blockquote>\n<ol>\n<li><p>Are the values in <code>thresholded prediction</code> some 0 or 1 (actually True or False?), how to get max-pooling on them? If I'm wrong, could you tell more about how to apply threshold and how to get chunk level prediction.</p></li>\n<li><p>In <code>Third Stage</code>, after you correcting the chunk-level label, is the question turn to strong-label-SED? or it's still a weak-label manner ?</p></li>\n</ol>",
          "rawMarkdown": "Thanks for replying.\nI read your solution again and have some questions. If you can talk more details that would be very nice.\n\n>After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction.\n\n1. Are the values in `thresholded prediction` some 0 or 1 (actually True or False?), how to get max-pooling on them? If I'm wrong, could you tell more about how to apply threshold and how to get chunk level prediction.\n\n2. In `Third Stage`, after you correcting the chunk-level label, is the question turn to strong-label-SED? or it's still a weak-label manner ?",
          "votes": 1
        },
        {
          "id": 1015106,
          "postDate": "2020-09-18T00:11:57.367Z",
          "content": "<blockquote>\n  <p>Are the values in thresholded prediction some 0 or 1 (actually True or False?), how to get max-pooling on them? If I'm wrong, could you tell more about how to apply threshold and how to get chunk level prediction.</p>\n</blockquote>\n<p>I updated the post, check out above :)</p>\n<blockquote>\n  <p>In Third Stage, after you correcting the chunk-level label, is the question turn to strong-label-SED? or it's still a weak-label manner ?</p>\n</blockquote>\n<p>Yes, I corrected the chunk-level label. I managed to use prediction as a label for strong-label-SED but it turned out the model that produced the prediction wasn't good enough, so I quit using it in strong-label manner and switched to use that to just correct the given weak label.</p>",
          "rawMarkdown": "> Are the values in thresholded prediction some 0 or 1 (actually True or False?), how to get max-pooling on them? If I'm wrong, could you tell more about how to apply threshold and how to get chunk level prediction.\n\nI updated the post, check out above :)\n\n> In Third Stage, after you correcting the chunk-level label, is the question turn to strong-label-SED? or it's still a weak-label manner ?\n\nYes, I corrected the chunk-level label. I managed to use prediction as a label for strong-label-SED but it turned out the model that produced the prediction wasn't good enough, so I quit using it in strong-label manner and switched to use that to just correct the given weak label.",
          "votes": 1
        },
        {
          "id": 1015237,
          "postDate": "2020-09-18T04:10:53.570Z",
          "content": "<p>Thanks, you're really awesome.</p>",
          "rawMarkdown": "Thanks, you're really awesome."
        }
      ]
    },
    {
      "id": 1012583,
      "postDate": "2020-09-16T07:02:57.413Z",
      "content": "<p>Congratulations. I was really impressed with what you shared during this competition I learnt a lot from your discussions and kernels.</p>",
      "rawMarkdown": "Congratulations. I was really impressed with what you shared during this competition I learnt a lot from your discussions and kernels.",
      "votes": 1,
      "replies": [
        {
          "id": 1013039,
          "postDate": "2020-09-16T13:23:26.993Z",
          "content": "<p>Thank you! and congratulations <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> for winning!</p>",
          "rawMarkdown": "Thank you! and congratulations @taggatle for winning!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1012540,
      "postDate": "2020-09-16T06:36:51.990Z",
      "content": "<p>Congrats and thanks for sharing details solution..</p>",
      "rawMarkdown": "Congrats and thanks for sharing details solution..",
      "votes": 1
    },
    {
      "id": 1012445,
      "postDate": "2020-09-16T05:21:30.977Z",
      "content": "<p>Congratulations, and thank your for sharing your work, It is really impressive. Good to know that label self-distillation is working here. I just started experimenting with it, but couldn't do much in the last 1-2 days before the competition end.</p>",
      "rawMarkdown": "Congratulations, and thank your for sharing your work, It is really impressive. Good to know that label self-distillation is working here. I just started experimenting with it, but couldn't do much in the last 1-2 days before the competition end.",
      "votes": 1,
      "replies": [
        {
          "id": 1013034,
          "postDate": "2020-09-16T13:20:39.250Z",
          "content": "<p>Thanks! <br>\nIn fact, I saw you sharing a lot in the PANDA competition and wanted to behave like you :)</p>\n<blockquote>\n  <p>Good to know that label self-distillation is working here.</p>\n</blockquote>\n<p>I now learned a new keyword: self-distillation, thanks!</p>",
          "rawMarkdown": "Thanks! \nIn fact, I saw you sharing a lot in the PANDA competition and wanted to behave like you :)\n\n> Good to know that label self-distillation is working here.\n\nI now learned a new keyword: self-distillation, thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1012424,
      "postDate": "2020-09-16T04:57:48.233Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>! I really appreciate your two notebooks and discussions. As a beginner for Kaggle competitions, I learn a lot from them. It is very kind of you to share these details. Thank you!</p>",
      "rawMarkdown": "Congrats @hidehisaarai1213! I really appreciate your two notebooks and discussions. As a beginner for Kaggle competitions, I learn a lot from them. It is very kind of you to share these details. Thank you!",
      "votes": 1
    },
    {
      "id": 1012263,
      "postDate": "2020-09-16T02:09:59.010Z",
      "content": "<p>Thank you sharing solution, and congrats!</p>\n<p>I had tried SED approach but it too late…<br>\nYour notebook was very useful and my work made easy!</p>\n<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you sharing solution, and congrats!\n\nI had tried SED approach but it too late...\nYour notebook was very useful and my work made easy!\n\nThank you very much!",
      "votes": 1,
      "replies": [
        {
          "id": 1013026,
          "postDate": "2020-09-16T13:16:36.840Z",
          "content": "<p>Thanks!<br>\nI also found you shared many interesting ideas, thanks for all those things 😊</p>",
          "rawMarkdown": "Thanks!\nI also found you shared many interesting ideas, thanks for all those things 😊"
        }
      ]
    },
    {
      "id": 1012223,
      "postDate": "2020-09-16T01:23:13.873Z",
      "content": "<p>Congrats solo gold medal! Your SED kernel has helped, thank you!</p>",
      "rawMarkdown": "Congrats solo gold medal! Your SED kernel has helped, thank you!",
      "votes": 1
    },
    {
      "id": 1012213,
      "postDate": "2020-09-16T01:11:09.830Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>  for your great work during all this compet ! I learnt a lot from  your various kernels !</p>",
      "rawMarkdown": "Congrats @hidehisaarai1213  for your great work during all this compet ! I learnt a lot from  your various kernels !",
      "votes": 1,
      "replies": [
        {
          "id": 1013024,
          "postDate": "2020-09-16T13:15:29.177Z",
          "content": "<p>Thanks! and congrats <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> for winning the 3rd place!<br>\nYou also shared a lot, thanks for all those things :)</p>",
          "rawMarkdown": "Thanks! and congrats @kneroma for winning the 3rd place!\nYou also shared a lot, thanks for all those things :)"
        }
      ]
    },
    {
      "id": 1012191,
      "postDate": "2020-09-16T00:48:46.183Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , this was my first audio competition, I stuck around throughout the competition and learnt a lot from your teachings, from your baseline notebook to SED notebook.</p>",
      "rawMarkdown": "Congratulations @hidehisaarai1213 , this was my first audio competition, I stuck around throughout the competition and learnt a lot from your teachings, from your baseline notebook to SED notebook.",
      "votes": 1
    },
    {
      "id": 1012175,
      "postDate": "2020-09-16T00:36:48.550Z",
      "content": "<p>Excellent work. For me, this is something incredible</p>",
      "rawMarkdown": "Excellent work. For me, this is something incredible",
      "votes": 1,
      "replies": [
        {
          "id": 1012217,
          "postDate": "2020-09-16T01:13:55.590Z",
          "content": "<p>Thanks! I used my brain a lot to come up with the ideas😎</p>",
          "rawMarkdown": "Thanks! I used my brain a lot to come up with the ideas😎"
        }
      ]
    },
    {
      "id": 1013092,
      "postDate": "2020-09-16T13:54:54.027Z",
      "content": "<p>great new concepts. :)</p>",
      "rawMarkdown": "great new concepts. :)",
      "votes": 2
    },
    {
      "id": 1012355,
      "postDate": "2020-09-16T03:35:27.643Z",
      "content": "<p>You are an inspiration, well done, SED seemed very useful to identify, I should have paid more attention.</p>",
      "rawMarkdown": "You are an inspiration, well done, SED seemed very useful to identify, I should have paid more attention.",
      "votes": 2
    },
    {
      "id": 1012169,
      "postDate": "2020-09-16T00:33:19.863Z",
      "content": "<p>Congrats and thanks for sharing details solution <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> </p>",
      "rawMarkdown": "Congrats and thanks for sharing details solution @hidehisaarai1213 ",
      "votes": 2,
      "replies": [
        {
          "id": 1012214,
          "postDate": "2020-09-16T01:11:21.290Z",
          "content": "<p>Thank you! I'll update it later to put some more detail and make it easier to read :)</p>",
          "rawMarkdown": "Thank you! I'll update it later to put some more detail and make it easier to read :)"
        }
      ]
    },
    {
      "id": 1013343,
      "postDate": "2020-09-16T16:46:00.073Z",
      "content": "<p>Nice solution. I learn very much from your notebooks. Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> </p>",
      "rawMarkdown": "Nice solution. I learn very much from your notebooks. Thanks @hidehisaarai1213 "
    },
    {
      "id": 1333325,
      "postDate": "2021-06-02T16:24:01.307Z",
      "content": "<p>Congrats on 6th place and thank you for your great solution.<br>\nThe walk of get gold medals while sharing a good notebooks isn't really easy.</p>\n<p>Thanks again and congrats on the gold medal in the 2021 competition as well.</p>",
      "rawMarkdown": "Congrats on 6th place and thank you for your great solution.\nThe walk of get gold medals while sharing a good notebooks isn't really easy.\n\nThanks again and congrats on the gold medal in the 2021 competition as well."
    },
    {
      "id": 1332268,
      "postDate": "2021-06-02T02:39:45.300Z",
      "content": "<p>I found this solution from <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243293\" target=\"_blank\">a BirdCLEF's top solution</a><br>\nThis post is so well-explained and educational. Will definitely come back and read in detail again</p>",
      "rawMarkdown": "I found this solution from [a BirdCLEF's top solution](https://www.kaggle.com/c/birdclef-2021/discussion/243293)\nThis post is so well-explained and educational. Will definitely come back and read in detail again"
    },
    {
      "id": 1021708,
      "postDate": "2020-09-22T04:57:24.370Z",
      "content": "<p>Thanks for sharing! learned a lot from your code.</p>",
      "rawMarkdown": "Thanks for sharing! learned a lot from your code."
    },
    {
      "id": 1020705,
      "postDate": "2020-09-21T11:19:40.050Z",
      "content": "<p>Amazing work , Congrats </p>",
      "rawMarkdown": "Amazing work , Congrats \n"
    },
    {
      "id": 1020533,
      "postDate": "2020-09-21T08:27:07.140Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!\n"
    },
    {
      "id": 1019841,
      "postDate": "2020-09-20T17:34:28.230Z",
      "content": "<p>Great work</p>",
      "rawMarkdown": "Great work"
    },
    {
      "id": 1019610,
      "postDate": "2020-09-20T14:36:22.937Z",
      "content": "<p>Great! Keep going!</p>",
      "rawMarkdown": "Great! Keep going!"
    },
    {
      "id": 1019609,
      "postDate": "2020-09-20T14:35:28.157Z",
      "content": "<p>good one there</p>",
      "rawMarkdown": "good one there"
    },
    {
      "id": 1018801,
      "postDate": "2020-09-20T01:45:46.300Z",
      "content": "<p>Great Work</p>",
      "rawMarkdown": "Great Work"
    },
    {
      "id": 1018436,
      "postDate": "2020-09-19T17:21:05.143Z",
      "content": "<p>Congrats!!</p>",
      "rawMarkdown": "Congrats!!"
    },
    {
      "id": 1018290,
      "postDate": "2020-09-19T15:24:50.240Z",
      "content": "<p>Nice work!!</p>",
      "rawMarkdown": "Nice work!!"
    },
    {
      "id": 1014270,
      "postDate": "2020-09-17T09:56:30.877Z",
      "content": "<p>wery well done :)</p>",
      "rawMarkdown": "wery well done :)"
    },
    {
      "id": 1012164,
      "postDate": "2020-09-16T00:30:10.157Z",
      "content": "<p>Thank you for sharing this. Very helpful!</p>",
      "rawMarkdown": "Thank you for sharing this. Very helpful!",
      "replies": [
        {
          "id": 1012212,
          "postDate": "2020-09-16T01:09:58.263Z",
          "content": "<p>Thank you for hosting such an interesting competition! I hope it helps your work :)</p>",
          "rawMarkdown": "Thank you for hosting such an interesting competition! I hope it helps your work :)"
        }
      ]
    },
    {
      "id": 1013820,
      "postDate": "2020-09-17T02:17:56.330Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1014250,
      "postDate": "2020-09-17T09:45:23.157Z",
      "content": "<p>Thanks so much :)</p>",
      "rawMarkdown": "Thanks so much :)",
      "votes": 1
    },
    {
      "id": 1014075,
      "postDate": "2020-09-17T07:33:03.313Z",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": 1
    },
    {
      "id": 1013890,
      "postDate": "2020-09-17T03:35:48.613Z",
      "content": "<p>thanks for easy explanation</p>",
      "rawMarkdown": "thanks for easy explanation",
      "votes": 1
    },
    {
      "id": 1013832,
      "postDate": "2020-09-17T02:46:58.040Z",
      "content": "<p>Thank you !!!</p>",
      "rawMarkdown": "Thank you !!!",
      "votes": 1
    },
    {
      "id": 1166156,
      "postDate": "2021-01-23T12:43:03.017Z",
      "content": "<p>Wow. This was educational. Thank you.</p>",
      "rawMarkdown": "Wow. This was educational. Thank you."
    },
    {
      "id": 1019017,
      "postDate": "2020-09-20T06:22:15.200Z",
      "content": "<p>Thank you !!</p>",
      "rawMarkdown": "Thank you !!"
    }
  ],
  "comments": [
    {
      "id": 1013294,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2020-09-16T16:10:07.970000",
      "content": "<p>Thanks for the extensive write-up and all of your engagement in the discussion forums! As you said, it's a hard task, and I don't think nearly as many people would have gotten a good start on it without your help.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1137457,
      "author_name": "Buffalo Spdwy",
      "author_url": "",
      "post_date": "2021-01-04T00:33:10.763000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , thanks for sharing! I am learning what you shared here for the new Rainforest competition. :)<br>\nBut I have a question about the 3 channel inputs <code>[normal logmel, PCEN, librosa.power_to_db(melspec ** 1.5)]</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F65df09a5f2d60e6e8e7889f1355bed18%2FScreenshot%20from%202021-01-03%2018-28-45.png?generation=1609720253136614&amp;alt=media\" alt=\"\"></p>\n<p><code>power_to_db</code> is taking a log. Raising to a power before taking the log is just scaling the output. Can it really help the model? </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1013417,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-09-16T17:26:54.877000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for all your selfless sharing. <br>\nYou deserve a Gold just for that !</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1012182,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-09-16T00:41:19.177000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> on the gold medal. Thanks so much for sharing a detailed solution and for all your contributions in this cmpetition.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1012177,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-09-16T00:37:40.920000",
      "content": "<p>Thanks a lot for all you've done for the competition, and congratz on the nice finish !</p>\n<p>Although I didn't believe in the SED idea (we didn't try it), we're glad that you made it work. I'll read your solution after a good night of sleep, giving it the full attention it deserves !</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1012221,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T01:20:43.117000",
          "content": "<p>Congratz again <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> !</p>\n<blockquote>\n  <p>Although I didn't believe in the SED idea (we didn't try it)</p>\n</blockquote>\n<p>Haha, I somehow knew this during the competition 😄</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1012174,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-09-16T00:36:48.643000",
      "content": "<p>Congrats again. I need to read this again in detail as it looks more elaborate than the noisy student method I wanted to try.</p>\n<p>I want to congratulate you for sharing so much as well.  </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1012219,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T01:16:45.120000",
          "content": "<p>Thank you!</p>\n<blockquote>\n  <p>the noisy student method</p>\n</blockquote>\n<p>I actually thought of using noisy student method but couldn't make it work well 😑</p>\n<blockquote>\n  <p>I want to congratulate you for sharing so much as well.</p>\n</blockquote>\n<p>Thanks! This word is also for you! Thanks for actively sharing lots of things on discussion!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1012239,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-16T01:34:41.733000",
          "content": "<blockquote>\n  <p>I actually thought of using noisy student method but couldn't make it work well</p>\n</blockquote>\n<p>Good to know, I have no regrets then :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1015686,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-09-18T11:04:17.820000",
      "content": "<p>Congratz and thank you for the summary.</p>\n<p>I was hoping for you all the best since the start of your competiton due to your great notebooks and explanations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1015685,
      "author_name": "Vimal Shetty",
      "author_url": "",
      "post_date": "2020-09-18T11:04:04.380000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1015408,
      "author_name": "Mikel Armendariz",
      "author_url": "",
      "post_date": "2020-09-18T07:04:25.307000",
      "content": "<p>Great solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1015366,
      "author_name": "Jaynil Gaglani",
      "author_url": "",
      "post_date": "2020-09-18T06:27:54.567000",
      "content": "<p>Thanks for writing this article. It really helps new folks like me out there!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1014280,
      "author_name": "Jerome Botie",
      "author_url": "",
      "post_date": "2020-09-17T10:07:56.173000",
      "content": "<p>That's what's up :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1014226,
      "author_name": "Subhabrata Mukherjee",
      "author_url": "",
      "post_date": "2020-09-17T09:17:40.747000",
      "content": "<p>It's awesome. Thanks for being helpful to all by sharing your works. A lot to learn!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1013274,
      "author_name": "Wessam Hamdy",
      "author_url": "",
      "post_date": "2020-09-16T15:46:04.070000",
      "content": "<p>Excellent good job</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1013120,
      "author_name": "Dave E",
      "author_url": "",
      "post_date": "2020-09-16T14:11:35.930000",
      "content": "<p>Congrats, great work and thanks so much for sharing during comp and now. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012777,
      "author_name": "R. Joseph Manoj, PhD",
      "author_url": "",
      "post_date": "2020-09-16T09:31:32.187000",
      "content": "<p>Excellent.. and great sharing. exposed some new concepts which will help me a lot. Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012655,
      "author_name": "LuisBlanche",
      "author_url": "",
      "post_date": "2020-09-16T07:58:45.023000",
      "content": "<p>You da real MVP <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> ! !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012592,
      "author_name": "Yu Kang",
      "author_url": "",
      "post_date": "2020-09-16T07:06:45.043000",
      "content": "<p>Congrats and excellent work!</p>\n<blockquote>\n  <p>and use EMA model for inference.</p>\n</blockquote>\n<p>Sorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1013036,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T13:22:18.037000",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Sorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.</p>\n</blockquote>\n<p>It's Exponential Moving Average in weight space. OK, I'll add some explanation :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1013902,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-09-17T03:49:28.030000",
          "content": "<p>Thanks for replying.<br>\nI read your solution again and have some questions. If you can talk more details that would be very nice.</p>\n<blockquote>\n  <p>After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction.</p>\n</blockquote>\n<ol>\n<li><p>Are the values in <code>thresholded prediction</code> some 0 or 1 (actually True or False?), how to get max-pooling on them? If I'm wrong, could you tell more about how to apply threshold and how to get chunk level prediction.</p></li>\n<li><p>In <code>Third Stage</code>, after you correcting the chunk-level label, is the question turn to strong-label-SED? or it's still a weak-label manner ?</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1015106,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-18T00:11:57.367000",
          "content": "<blockquote>\n  <p>Are the values in thresholded prediction some 0 or 1 (actually True or False?), how to get max-pooling on them? If I'm wrong, could you tell more about how to apply threshold and how to get chunk level prediction.</p>\n</blockquote>\n<p>I updated the post, check out above :)</p>\n<blockquote>\n  <p>In Third Stage, after you correcting the chunk-level label, is the question turn to strong-label-SED? or it's still a weak-label manner ?</p>\n</blockquote>\n<p>Yes, I corrected the chunk-level label. I managed to use prediction as a label for strong-label-SED but it turned out the model that produced the prediction wasn't good enough, so I quit using it in strong-label manner and switched to use that to just correct the given weak label.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1015237,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-09-18T04:10:53.570000",
          "content": "<p>Thanks, you're really awesome.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1012583,
      "author_name": "Ryan Wong",
      "author_url": "",
      "post_date": "2020-09-16T07:02:57.413000",
      "content": "<p>Congratulations. I was really impressed with what you shared during this competition I learnt a lot from your discussions and kernels.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1013039,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T13:23:26.993000",
          "content": "<p>Thank you! and congratulations <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> for winning!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1012540,
      "author_name": "Nehal Alex",
      "author_url": "",
      "post_date": "2020-09-16T06:36:51.990000",
      "content": "<p>Congrats and thanks for sharing details solution..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012445,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-09-16T05:21:30.977000",
      "content": "<p>Congratulations, and thank your for sharing your work, It is really impressive. Good to know that label self-distillation is working here. I just started experimenting with it, but couldn't do much in the last 1-2 days before the competition end.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1013034,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T13:20:39.250000",
          "content": "<p>Thanks! <br>\nIn fact, I saw you sharing a lot in the PANDA competition and wanted to behave like you :)</p>\n<blockquote>\n  <p>Good to know that label self-distillation is working here.</p>\n</blockquote>\n<p>I now learned a new keyword: self-distillation, thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1012424,
      "author_name": "Po-Kuan Wu",
      "author_url": "",
      "post_date": "2020-09-16T04:57:48.233000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>! I really appreciate your two notebooks and discussions. As a beginner for Kaggle competitions, I learn a lot from them. It is very kind of you to share these details. Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012263,
      "author_name": "Takamichi Toda",
      "author_url": "",
      "post_date": "2020-09-16T02:09:59.010000",
      "content": "<p>Thank you sharing solution, and congrats!</p>\n<p>I had tried SED approach but it too late…<br>\nYour notebook was very useful and my work made easy!</p>\n<p>Thank you very much!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1013026,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T13:16:36.840000",
          "content": "<p>Thanks!<br>\nI also found you shared many interesting ideas, thanks for all those things 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1012223,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2020-09-16T01:23:13.873000",
      "content": "<p>Congrats solo gold medal! Your SED kernel has helped, thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012213,
      "author_name": "kkiller",
      "author_url": "",
      "post_date": "2020-09-16T01:11:09.830000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>  for your great work during all this compet ! I learnt a lot from  your various kernels !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1013024,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T13:15:29.177000",
          "content": "<p>Thanks! and congrats <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> for winning the 3rd place!<br>\nYou also shared a lot, thanks for all those things :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1012191,
      "author_name": "Alan Choon",
      "author_url": "",
      "post_date": "2020-09-16T00:48:46.183000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , this was my first audio competition, I stuck around throughout the competition and learnt a lot from your teachings, from your baseline notebook to SED notebook.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1012175,
      "author_name": "ermacov",
      "author_url": "",
      "post_date": "2020-09-16T00:36:48.550000",
      "content": "<p>Excellent work. For me, this is something incredible</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1012217,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T01:13:55.590000",
          "content": "<p>Thanks! I used my brain a lot to come up with the ideas😎</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1013092,
      "author_name": " Akash Gupta",
      "author_url": "",
      "post_date": "2020-09-16T13:54:54.027000",
      "content": "<p>great new concepts. :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1012355,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2020-09-16T03:35:27.643000",
      "content": "<p>You are an inspiration, well done, SED seemed very useful to identify, I should have paid more attention.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1012169,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-09-16T00:33:19.863000",
      "content": "<p>Congrats and thanks for sharing details solution <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1012214,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-09-16T01:11:21.290000",
          "content": "<p>Thank you! I'll update it later to put some more detail and make it easier to read :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1013343,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-16T16:46:00.073000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1333325,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-02T16:24:01.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1332268,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-02T02:39:45.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1021708,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-22T04:57:24.370000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1020705,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-21T11:19:40.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1020533,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-21T08:27:07.140000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1019841,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-20T17:34:28.230000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1019610,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-20T14:36:22.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1019609,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-20T14:35:28.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1018801,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-20T01:45:46.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1018436,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-19T17:21:05.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1018290,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-19T15:24:50.240000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1014270,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T09:56:30.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1012164,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-16T00:30:10.157000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1012212,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-16T01:09:58.263000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1013820,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T02:17:56.330000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1014250,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T09:45:23.157000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1014075,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T07:33:03.313000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1013890,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T03:35:48.613000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1013832,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T02:46:58.040000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1166156,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-23T12:43:03.017000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1019017,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-20T06:22:15.200000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012159": "First of all, I would like to sincerely thank [@stefankahl](https://www.kaggle.com/stefankahl), [@tomdenton](https://www.kaggle.com/tomdenton), [@holgerklinck](https://www.kaggle.com/holgerklinck), and the members of kaggle team for hosting this competition. I had a lot of fun tackling on some of the challenging problems of machine learning thinking of the generative process of the data. Also many thanks to [@hengck23](https://www.kaggle.com/hengck23) for actively sharing a lot of deep insights. It helped me a lot to come up with some good ideas and also made me convinced that I was in good direction.\n\nFollowing the recent two competitions: [PANDA](https://www.kaggle.com/c/prostate-cancer-grade-assessment) challenge and [GWD](https://www.kaggle.com/c/global-wheat-detection) challenge, this competition was also about **domain shift** and **noisy labels**.\nCombination of these two challenging topics made this competition extremely difficult and we were troubled a lot how to make stable validation scheme. To be honest, contrary to [@cpmpml](https://www.kaggle.com/cpmpml)'s [expectation](https://www.kaggle.com/c/birdsong-recognition/discussion/181499#1010570), I couldn't find any good local validation scheme as the labels of training dataset contains a lot of noise. `Trust LB` was also not a very good policy since public LB was only 27% and we didn't know how the test set was devided. Instead, I took the policy of [PANDA's competitors](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230) - **Ignore CV, care about public LB, and trust methodology**.\n\nAlthough this competition was difficult as a data science competition, it was actually pretty close to a realistic setting and was full of problems that we often face in applying data science to real-world problem. Especially the combination of domain shift and noisy labels often happens (I think) when we are to use data from User Generated Contents(UGC) web service like Xeno Canto, YouTube, Twitter for training machine learning algorithms. Therefore, I think my solution is useful not only for this competition but also for those data science tasks related with UGC data, as it's basically focused on dealing with noisy labels and domain shift.\n\n## Solution in three lines\n\n* 3 stages of training to gradually remove noise in labels\n* SED style training and inference as I introduced [here](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)\n* Ensemble of 11 EMA models trained with the whole dataset / whole [extended dataset](https://www.kaggle.com/c/birdsong-recognition/discussion/159970) to stabilize the result\n\nCode is available here: [koukyo1994/kaggle-birdcall-6th-place](https://github.com/koukyo1994/kaggle-birdcall-6th-place).\nNote: I've re-implemented the code from original repository since it's quite messy. However, I haven't checked whether whole pipeline works; if you find something wrong with the code, please let me know :)\n\n## Motivation\n\nIt's quite obvious that there is a huge gap between training dataset and test dataset: it was described by the host and can also be seen through submission. However, *how* they are different is not obvious. Therefore, the first thing I did was to understand by observation what kind of domain shifts are seen between training and test datasets.\n\nDomain shift is an umbrella term and there are several problem classes of domain shift. The most famous one is *covariate shift*, where \\\\( P_{train}(X)\\neq P_{test}(X) \\\\) and \\\\( P_{train}(Y|X)=P_{test}(Y|X) \\\\). This one is well-studied and several algorithms are proposed to deal with the situation, but is not the type of domain shift in this competition. Another problem class is *prior probability shift* or *target shift*, where \\\\( P_{train}(Y)\\neq P_{test}(Y) \\\\) and \\\\( P_{train}(X|Y)=P_{test}(X|Y) \\\\), also not the type in this competition because \\\\( P_{train}(X|Y) \\\\) is not the same as \\\\( P_{test}(X|Y) \\\\). In fact, in this competition, multiple distribution shifts are present - shift in input space, shift in prior probability of labels, and shift in the function which connects \\\\( X \\\\) and \\\\( Y \\\\).\n\nHow should we tackle a problem with various distribution shifts? The answer is simple - *divide the difficulty*. As I wrote above, shifts were three folds:\n\n1) shift in input space. For example, SNR difference or difference in sound collection environment (device/sampling rate/temperature/weather/...) between train and test. Occurence of non-target sound events is also a part of this shift.\n2) shift in prior probability of labels. Distribution difference of species or distribution difference of calltypes, or else.\n3) shift in the function which connects \\\\( X \\\\) and \\\\( Y \\\\). This has a very strong relation with label noise. Thinking of how the labels were created in train dataset and in test dataset, one could come up with the fact that Label Function(LF) of train dataset is completely different from that of test dataset. The former is annotations of the uploader (and can have large variation), whereas the latter is probably those of dedicated annotator(s) (and possibly have smaller variation).\n\nI decided to address these one by one and applied the following techniques.\n\n* For 1), providing all the possible variation for train dataset may help. This is done by data augmentation.\n* For 2), I just couldn't come up with smart ideas. I used ensemble of multiple models trained with datasets with different distributions of the labels to address this, but I think that is suboptimal.\n* For 3), correcting the labels of train dataset to make train LF closer to test LF can help.\n\nOn the other hand, label noise is also a term that contains multiple problem classes. First, in this competition, labels of train dataset are provided as *weak labels*. As @hengck23 pointed out [here](https://www.kaggle.com/c/birdsong-recognition/discussion/174774#972122), weak label can be treated as noisy label if we change the point of view. Also there are some missing labels, which I'll explain later.\n\nI also decided to address these one by one.\n\n* For weak label as noisy label, I first train a model with long chunk and then use the prediction as corrected label. I at first tried to create strong labels but couldn't make the first stage model good enough, therefore I used the prediction to correct weak labels we have in train dataset.\n* For missing label, I also used the prediction of a model to find those.\n\n## First stage - build a model useful enough for addressing missing labels\n\nIn this stage, I used PANNs model. I used some basic augmentations (`NoiseInjection`, `PitchShift`, `RandomVolume`) and used `secondary_labels`. The keys in this stage were two folds:\n\n* train with long chunk(30s) so that it would include call events of the species in `primaly_label` and `secondary_labels`\n* use attention pooling and max pooling to get weak prediction from `framewise_output`\n\nHere are the reason behind.\n\n### train with long chunk\n\nAssume we have `primary_label` of `birdA` and `secondary_labels` of `birdB` and `birdC`. Melspectrogram of the corresponding audio clip is something like the figure below (sorry for my poor drawing). If we use small window size, it may not include any sound events or include some sound events but not enough for the given labels. To make the model learn correctly, we need to make each label correspond to call event(s) of each species. For this reason, I used long chunk. Maybe it is better to use longer chunk like 1 minutes or more but I compromised to use 30s chunk considering the time for computation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1489975%2Fcd841cb51b7b1fa4de6a6b2b7880e2d6%2FIMG_BBFA97950653-1.jpeg?generation=1600216062555033&alt=media)\n\n### Combination of attention pooling and max pooling to get weak prediction\n\nThis is something I shared [here](https://www.kaggle.com/c/birdsong-recognition/discussion/167611).\nIn the comments in [my SED notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection), some said there were lots of false positives. This comes from the following reason.\n\nWeak prediction from attention pooling is made by applying self-attention filter on `framewise_outputs`. Now, we only have weak labels, we calculate the loss with weak predictions and weak labels. The gradient will be distributed to self-attention layer and pointwise classifier, but for pointwise classifier, strong supervision comes when self-attention put high probability at that point. Therefore, pointwise classifier are more likely to produce high probability value and rely on the attention layer to reduce false positives. This is not good when we are interested in the output of pointwise classifier(`framewise_outputs`).\n\nOn the contrary, max pooling suppresses high probability values come out from pointwise classifier but it also has a defect that it is weak to impulse noise. Therefore, combination of max pooling and attention pooling can make balanced prediction and we can expect that be good prediction. For this reason, I used both but not combining the weak prediction from each aggregation but use each output and calculate loss for each, and sum them up. Therefore, the loss function I used in this stage was like this\n\n```python\nbce = BCELoss()\nloss_att = bce(weak_pred_with_attention, label)\nloss_max = bce(weak_pred_with_maxpooling, label)\nloss = 1.0 * loss_att + 0.5 * loss_max\n```\n\n### Summarize this stage\n\n* Single PANNs model\n* BCE on `clipwise_output` and also on maxpooled output.\n* Adam + CosineAnnealing, 55epochs training\n* train with randomly cropped 30s chunk\n* validate on randomly cropped 30s chunk\n* Augmentations on raw waveform\n  - `NoiseInjection` (max noise amplitude 0.04)\n  - `PitchShift` (max pitch level 3)\n  - `RandomVolume` (max db level 4)\n\n### Get oof prediction and use it to find missing labels\n\nWith training procedure above, the model would get around 0.575 - 0.578. In fact, this is the weight I used in the public notebook.\nI trained 5folds and got oof prediction on the whole training set. Then I used this oof prediction to find missing labels.\n\nMissing labels are more likely to be found from samples that does not have `secondary_labels`. It is up to the uploader to fill in `secondary_labels` or `background`, so some uploaders may not feel like to fill in those. Therefore, I picked samples without `secondary_labels` and used oof prediction of those to get additional labels if the probability of species that are not in their `primary_label` is over 0.9.\n\n## Second stage - build a model with additional labels to get stronger labels\n\nIn this stage, I used SED model with ResNeSt encoder. The difference between first stage and second stage is not that big - only the model, the existence of found labels (the labels obtained from the oof prediction of the first stage), and the input. I started to use 3channel input. The first channel was normal log-melspectrogram and the second channel was PCEN. The third channel was also log-melspectrogram but instead of using `librosa.power_to_db(melspec)`, I used `librosa.power_to_db(melspec ** 1.5)`. The idea of using different input for each channel comes from [this post](https://www.kaggle.com/c/birdsong-recognition/discussion/170959). The chunk size is also reduced to 20s because of the GPU memory size limitation.\n\nI also changed the attention pooling slightly given the [advice of @hengck23](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments) to use `torch.tanh` instead of `torch.clamp`.\n\nThe result of this stage got around 0.60. I also trained 5folds model and this time I got oof prediction of `framewise_outputs`.\n\nTo summarize,\n\n* SED model with ResNeSt50 encoder, attention pooling head of PANNs (`torch.clamp` -> `torch.tanh`)\n* BCE on `clipwise_output` and also on maxpooled output\n* Adam + CosineAnnealing, 75epochs training\n* Add additional `secondary_labels` found in stage1\n* 3channels input - \\[normal logmel, PCEN, `librosa.power_to_db(melspec ** 1.5)`\\]\n* train with randomly cropped 20s chunk\n* validate on randomly cropped 30s chunk\n* Augmentations on raw waveform\n  - `NoiseInjection` (max noise amplitude 0.04)\n  - `PitchShift` (max pitch level 3)\n  - `RandomVolume` (max db level 4)\n\n## Third stage - build a model with oof framewise_outputs\n\nDespite the use of training with long chunk size and missing labels, label noise problem is far from solved. SNR level to decide whether a call event was present is different between training set and test set. In fact, that was also different between samples of training set because basically each annotator (uploader) had their own labeling criteria. For this reason, I used the oof prediction of `framewise_outputs` of the second stage to further correct the labels of the training dataset. \n\nIn this stage, I also cropped 20s chunk randomly to get a batch. At this time I also cropped the corresponding part of the predicted `framewise_outputs` and apply threshold on them (threshold varies from 0.3 - 0.7). After thresholding, I get max pooling of the thresholded prediction in time axis to get chunk level prediction. Still, this prediction can be noisy and may contain false positives, I apply `logical_and` between predicted labels and provided labels. In this way, I got corrected chunk level label and use that for training.\n\nHere's the pseudo-code of the process above\n\n```python\ny_batch = y[start_index:end_index]\nsoft_label_batch = soft_label[start_index_for_label:end_index_for_label]  # (n_frames, 264)\nthresholded = (soft_label_batch >= threshold).astype(int)\n\nweak_pred = thresholded.max(axis=0)  # (264,)\ncorrected_label = np.logical_and(label, weak_pred)  # (264,)\n```\n\nIn this stage, I also tried EfficientNet-B0 encoder and FocalLoss. Combination of ResNeSt encoder and FocalLoss didn't work well, whereas EffNet-B0 and FocalLoss worked well on public LB.\n\nAll the other settings were the same as that of second stage. After this, I got around 0.61x score on public LB.\n\nTo summarize,\n\n* SED model with ResNeSt50 encoder, attention pooling head of PANNs (`torch.clamp` -> `torch.tanh`) / EfficientNet-B0 encoder\n* BCE on `clipwise_output` and also on maxpooled output / FocalLoss on `clipwise_output` and also on maxpooled output for EfficientNet-B0\n* Adam + CosineAnnealing, 75epochs training\n* Add additional `secondary_labels` found in stage1\n* Correct labels using the prediction of stage2 model.\n* 3channels input - \\[normal logmel, PCEN, `librosa.power_to_db(melspec ** 1.5)`\\]\n* train with randomly cropped 20s chunk\n* validate on randomly cropped 30s chunk\n* Augmentations on raw waveform\n  - `NoiseInjection` (max noise amplitude 0.04)\n  - `PitchShift` (max pitch level 3)\n  - `RandomVolume` (max db level 4)\n\n## Ensemble\n\nWith the corrected chunk level label, I trained the model with the whole dataset and use EMA model (using the implementation [here](https://pytorch.org/docs/stable/optim.html#stochastic-weight-averaging)) for inference. This is a technique also used in [GWD competition](https://www.kaggle.com/c/global-wheat-detection/discussion/172458). \n\nI prepared models with different threshold (threshold on `framewise_outputs` of oof prediction) to make the model robust. Also important was the use of extended dataset. It doesn't get better result on public LB but when I used that for ensemble, it bumped up the score.\n\n## Things that didn't work for me\n\n* Mixup\n* Calltype classification (781 class)\n* noisy student training\n* Larger models of efficientnet (b1, b2, b3...)\n* etc...\n\n## Things that worked but doesn't make sense to me\n\nWhen I change the PERIOD used for inference, the result greatly changed. Basically the larger value PERIOD is, the better the result. I just couldn't figure out why.\n\n## Things that I wanted to try but couldn't\n\nThere are lots of things I couldn't do due to the time limitation \n\n* Use a variety of models\n* Further refinement of the labels of train dataset/train extended dataset\n* Use of auxiliary predictor to predict longitude/latitude/elevation and use the prediction to correct the main classifier\n* Post-processing to refine prediction using species correlation information (I couldn't get the API key for ebird.org therefore I couldn't collect correlation information)\n* Mixing background noise\n* etc...\n\n## The thing that helped me a lot during competition\n\nI created a simple streamlit app to check the audio data. I mainly used this to check the effect of augmentations or to check the quality of SED models prediction. This helped me a lot to figure out the major problems of this competition\n\nhttps://github.com/koukyo1994/streamlit-audio\n\nLater I learned @fkubota also made an app with similar functionality. I didn't use this but it seems it's better than mine.\n\nhttps://github.com/fkubota/spectrogram-tree",
    "1013294": "Thanks for the extensive write-up and all of your engagement in the discussion forums! As you said, it's a hard task, and I don't think nearly as many people would have gotten a good start on it without your help.",
    "1137457": "Hi, @hidehisaarai1213 , thanks for sharing! I am learning what you shared here for the new Rainforest competition. :)\nBut I have a question about the 3 channel inputs ` [normal logmel, PCEN, librosa.power_to_db(melspec ** 1.5)]`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F65df09a5f2d60e6e8e7889f1355bed18%2FScreenshot%20from%202021-01-03%2018-28-45.png?generation=1609720253136614&alt=media)\n\n`power_to_db` is taking a log. Raising to a power before taking the log is just scaling the output. Can it really help the model? ",
    "1013417": "Thank you @hidehisaarai1213 for all your selfless sharing. \nYou deserve a Gold just for that !",
    "1012182": "Congrats @hidehisaarai1213 on the gold medal. Thanks so much for sharing a detailed solution and for all your contributions in this cmpetition.",
    "1012177": "Thanks a lot for all you've done for the competition, and congratz on the nice finish !\n\nAlthough I didn't believe in the SED idea (we didn't try it), we're glad that you made it work. I'll read your solution after a good night of sleep, giving it the full attention it deserves !",
    "1012174": "Congrats again. I need to read this again in detail as it looks more elaborate than the noisy student method I wanted to try.\n\nI want to congratulate you for sharing so much as well.  ",
    "1015686": "Congratz and thank you for the summary.\n\nI was hoping for you all the best since the start of your competiton due to your great notebooks and explanations!",
    "1015685": "Great work!",
    "1015408": "Great solution!",
    "1015366": "Thanks for writing this article. It really helps new folks like me out there!",
    "1014280": "That's what's up :)",
    "1014226": "It's awesome. Thanks for being helpful to all by sharing your works. A lot to learn!",
    "1013274": "Excellent good job",
    "1013120": "Congrats, great work and thanks so much for sharing during comp and now. ",
    "1012777": "Excellent.. and great sharing. exposed some new concepts which will help me a lot. Thanks",
    "1012655": "You da real MVP @hidehisaarai1213 ! !",
    "1012592": "Congrats and excellent work!\n\n> and use EMA model for inference.\n\nSorry but what is EMA, in the link of GWD comp also didn't talk the detail about it.",
    "1012583": "Congratulations. I was really impressed with what you shared during this competition I learnt a lot from your discussions and kernels.",
    "1012540": "Congrats and thanks for sharing details solution..",
    "1012445": "Congratulations, and thank your for sharing your work, It is really impressive. Good to know that label self-distillation is working here. I just started experimenting with it, but couldn't do much in the last 1-2 days before the competition end.",
    "1012424": "Congrats @hidehisaarai1213! I really appreciate your two notebooks and discussions. As a beginner for Kaggle competitions, I learn a lot from them. It is very kind of you to share these details. Thank you!",
    "1012263": "Thank you sharing solution, and congrats!\n\nI had tried SED approach but it too late...\nYour notebook was very useful and my work made easy!\n\nThank you very much!",
    "1012223": "Congrats solo gold medal! Your SED kernel has helped, thank you!",
    "1012213": "Congrats @hidehisaarai1213  for your great work during all this compet ! I learnt a lot from  your various kernels !",
    "1012191": "Congratulations @hidehisaarai1213 , this was my first audio competition, I stuck around throughout the competition and learnt a lot from your teachings, from your baseline notebook to SED notebook.",
    "1012175": "Excellent work. For me, this is something incredible",
    "1013092": "great new concepts. :)",
    "1012355": "You are an inspiration, well done, SED seemed very useful to identify, I should have paid more attention.",
    "1012169": "Congrats and thanks for sharing details solution @hidehisaarai1213 ",
    "1013343": "Nice solution. I learn very much from your notebooks. Thanks @hidehisaarai1213 ",
    "1333325": "Congrats on 6th place and thank you for your great solution.\nThe walk of get gold medals while sharing a good notebooks isn't really easy.\n\nThanks again and congrats on the gold medal in the 2021 competition as well.",
    "1332268": "I found this solution from [a BirdCLEF's top solution](https://www.kaggle.com/c/birdclef-2021/discussion/243293)\nThis post is so well-explained and educational. Will definitely come back and read in detail again",
    "1021708": "Thanks for sharing! learned a lot from your code.",
    "1020705": "Amazing work , Congrats \n",
    "1020533": "Great Work!\n",
    "1019841": "Great work",
    "1019610": "Great! Keep going!",
    "1019609": "good one there",
    "1018801": "Great Work",
    "1018436": "Congrats!!",
    "1018290": "Nice work!!",
    "1014270": "wery well done :)",
    "1012164": "Thank you for sharing this. Very helpful!",
    "1013820": "",
    "1014250": "Thanks so much :)",
    "1014075": "Thanks for sharing :)",
    "1013890": "thanks for easy explanation",
    "1013832": "Thank you !!!",
    "1166156": "Wow. This was educational. Thank you.",
    "1019017": "Thank you !!"
  }
}