{
  "id": 183571,
  "title": "7th place solution (writeup with code)",
  "url": "/competitions/birdsong-recognition/writeups/cp-jku-three-geese-and-a-gan-7th-place-solution-wr",
  "author_name": "",
  "post_date": "2020-12-21T16:58:59.300Z",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, congratulations to all the participants, and thanks a lot to the organizers for this competition! It was great fun!</p>\n<p>In short, our approach was based on <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">Jan's Sound Event Detection model from BirdCLEF 2018</a>, replacing the predictor with a pretrained CNN from <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">Qiuqiang Kong's PANN repository</a>. It was trained on 30-second snippets from the official training set extended with <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">Rohan Rao's xeno-canto crawls</a>, using binary cross-entropy against all labeled foreground and background species. Data was augmented by choosing random weights for downmixing the left and right channel to mono (for stereo files), and by mixing in bird-free background noise from the <a href=\"https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591\" target=\"_blank\">Chernobyl BiVA</a> and <a href=\"https://zenodo.org/record/1205569\" target=\"_blank\">BirdVox-full-night</a> datasets. Inference followed a two-stage procedure that first established a set of species for the recording, using 20-second windows and a threshold of 0.5, then looked for these species in 5-second windows with a threshold of 0.3. Code is <a href=\"https://github.com/f0k/kagglebirds2020\" target=\"_blank\">available on github</a>.</p>\n<p>The following sections will explain things in more detail, spread over additional posts due to Kaggle's post size limit. Choose \"Sort by: Oldest\" to see them in their original order.</p>\n<p>If you prefer a slide deck over text, feel free to click through <a href=\"https://docs.google.com/presentation/d/10m0W13sJozYmfWPcPlaBg1j-MOW7p6w_-Tp3pG5KGE0/edit?usp=sharing\" target=\"_blank\">some slides explaining what we did</a>.</p>",
  "messages": [
    {
      "id": "1014127",
      "postDate": "09/17/2020 08:22:59",
      "content": "<p>First of all, congratulations to all the participants, and thanks a lot to the organizers for this competition! It was great fun!</p>\n<p>In short, our approach was based on <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">Jan's Sound Event Detection model from BirdCLEF 2018</a>, replacing the predictor with a pretrained CNN from <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">Qiuqiang Kong's PANN repository</a>. It was trained on 30-second snippets from the official training set extended with <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">Rohan Rao's xeno-canto crawls</a>, using binary cross-entropy against all labeled foreground and background species. Data was augmented by choosing random weights for downmixing the left and right channel to mono (for stereo files), and by mixing in bird-free background noise from the <a href=\"https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591\" target=\"_blank\">Chernobyl BiVA</a> and <a href=\"https://zenodo.org/record/1205569\" target=\"_blank\">BirdVox-full-night</a> datasets. Inference followed a two-stage procedure that first established a set of species for the recording, using 20-second windows and a threshold of 0.5, then looked for these species in 5-second windows with a threshold of 0.3. Code is <a href=\"https://github.com/f0k/kagglebirds2020\" target=\"_blank\">available on github</a>.</p>\n<p>The following sections will explain things in more detail, spread over additional posts due to Kaggle's post size limit. Choose \"Sort by: Oldest\" to see them in their original order.</p>\n<p>If you prefer a slide deck over text, feel free to click through <a href=\"https://docs.google.com/presentation/d/10m0W13sJozYmfWPcPlaBg1j-MOW7p6w_-Tp3pG5KGE0/edit?usp=sharing\" target=\"_blank\">some slides explaining what we did</a>.</p>",
      "rawMarkdown": "First of all, congratulations to all the participants, and thanks a lot to the organizers for this competition! It was great fun!\n\nIn short, our approach was based on [Jan's Sound Event Detection model from BirdCLEF 2018](https://github.com/f0k/birdclef2018), replacing the predictor with a pretrained CNN from [Qiuqiang Kong's PANN repository](https://github.com/qiuqiangkong/audioset_tagging_cnn). It was trained on 30-second snippets from the official training set extended with [Rohan Rao's xeno-canto crawls](https://www.kaggle.com/c/birdsong-recognition/discussion/159970), using binary cross-entropy against all labeled foreground and background species. Data was augmented by choosing random weights for downmixing the left and right channel to mono (for stereo files), and by mixing in bird-free background noise from the [Chernobyl BiVA](https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591) and [BirdVox-full-night](https://zenodo.org/record/1205569) datasets. Inference followed a two-stage procedure that first established a set of species for the recording, using 20-second windows and a threshold of 0.5, then looked for these species in 5-second windows with a threshold of 0.3. Code is [available on github](https://github.com/f0k/kagglebirds2020).\n\nThe following sections will explain things in more detail, spread over additional posts due to Kaggle's post size limit. Choose \"Sort by: Oldest\" to see them in their original order.\n\nIf you prefer a slide deck over text, feel free to click through [some slides explaining what we did](https://docs.google.com/presentation/d/10m0W13sJozYmfWPcPlaBg1j-MOW7p6w_-Tp3pG5KGE0/edit?usp=sharing).",
      "votes": null
    },
    {
      "id": "1014359",
      "postDate": "09/17/2020 11:31:27",
      "content": "<blockquote>\n  <p>Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions</p>\n</blockquote>\n<p>Unfortunately this is true for many teams in every competition.  Looking for your writeup, and congrats on the result!</p>",
      "rawMarkdown": "> Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions\n\nUnfortunately this is true for many teams in every competition.  Looking for your writeup, and congrats on the result!",
      "votes": null
    },
    {
      "id": "1014384",
      "postDate": "09/17/2020 11:59:22",
      "content": "<p>Congratulations 🎊 </p>\n<p>Hope for next time…. Waiting for your writeup </p>",
      "rawMarkdown": "Congratulations 🎊 \n\nHope for next time.... Waiting for your writeup",
      "votes": null
    },
    {
      "id": "1014659",
      "postDate": "09/17/2020 16:19:19",
      "content": "<p>Congratulations!!!</p>",
      "rawMarkdown": "Congratulations!!!",
      "votes": null
    },
    {
      "id": "1121403",
      "postDate": "12/21/2020 16:08:54",
      "content": "<h2>Team</h2>\n<p>We are three PhD students (Khaled Koutini, Lukas Martak, Paul Primus) and a postdoc (Jan Schlüter) from the <a href=\"https://https://www.jku.at/en/institute-of-computational-perception/\" target=\"_blank\">Institute of Computational Perception</a> at Johannes-Kepler-University Linz, Austria. Since 20 of our 23 submissions were done by me (Jan), most of this is written from my perspective, but each of us did important contributions I will highlight below.</p>\n<h2>Network architecture</h2>\n<p>The general model architecture looks as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2Faf6861f1136591254f287a776e3682f1%2Fwriteup_network.png?generation=1608311897835230&amp;alt=media\" alt=\"\"></p>\n<p>From an (arbitrarily long) monophonic raw audio recording, the part denoted as \"Frontend\" computes a spectrogram-like representation. In the next step, a Fully-Convolutional Network (FCN) processes this representation into a time series of logits for every class. When passed through a sigmoid, these would give us local predictions at every time step. Since we do not have local labels to train these, only file-wise labels, we apply a global pooling operation (over time) to obtain a single logit per class. Passed through a sigmoid, these serve as our file-wise predictions.</p>\n<p>This layout matches what we used for <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">BirdCLEF 2018</a>, except that we are using a sigmoid instead of a softmax now (in BirdCLEF, evaluation was based on the ranking of species probabilities, for which a softmax helped). It also matches <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">what Hidehisa Arai explained</a> early on in the competition.</p>\n<p>We still have several options for the three components: The frontend, the local predictor, and the global pooling operation.</p>\n<h3>Frontend</h3>\n<p><br>\nPyTorch architecture listing:</p>\n<pre><code>(frontend): Sequential(\n  (filterbank): Sequential(\n    (stft): STFT(winsize=1024, hopsize=315, complex=False)\n    (melfilter): MelFilter(num_bands=80, min_freq=27.5, max_freq=10000.0)\n  )\n  (magscale): Log1p(trainable=True)\n  (denoise): SubtractMedian()\n  (norm): TemporalBatchNorm(\n    (bn): BatchNorm1d(80, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n  )\n)\n</code></pre>\n<p></p>\n<p>The frontend takes in audio recordings at a sample rate of 22050 Hz. (While the test recordings were at 32000 Hz, not all training recordings go that high, so I resampled everything to 22050 Hz.)</p>\n<p>It consists of:</p>\n<ul>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L171\" target=\"_blank\">STFT</a>:<br>\nwindow size of 1024 samples, hop size of 315 samples (resulting in 22050/315=70 frames per second), with Hann window, keeping only the magnitudes.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L268\" target=\"_blank\">mel filterbank</a>:<br>\n80 triangular filters from 27.5 Hz to 10 kHz (not going up to 11025 Hz on purpose to leave some room for pitch shifting)</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L72\" target=\"_blank\">nonlinear magnitude scaling</a>:<br>\nMagnitudes are compressed by passing them through \\(y = \\log(1 + 10^a x)\\), where \\(a\\) is initialized to zero and learned by backpropagation.<br>\nI got similar, but maybe slightly worse results with \\(y = x^{\\sigma(a)}\\), where \\(a\\) is initialized to zero (resulting in \\(y = \\sqrt{x}\\)) and learned by backpropagation.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L163\" target=\"_blank\">denoising</a>:<br>\nThe different recordings have very different background noise floors, both due to the different environments and due to different recording equipment. Human listeners are quite good at adapting to this within a few seconds. To make it easier for the network, I unify the recordings somewhat by subtracting the median over time from each frequency band (separately for each recording or excerpt). I also tried <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L114\" target=\"_blank\">Per-Channel Energy Normalization</a>, but subtracting the median worked better (and is faster). For PANNs this step is skipped, as it wouldn't match the input they are pretrained on.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L53\" target=\"_blank\">normalization</a>:<br>\nTo ensure inputs are in a reasonable range (and stay in a reasonable range when the magnitude scaling changes during training), each frequency band is normalized over time with Batch Normalization.</li>\n</ul>\n<h3>Local predictor</h3>\n<p>The purpose of the local predictor is to take the spectrogram produced by the frontend, and produce 200 time series of logits, one for each bird species.<br>\nThe spectrogram can be regarded as an 80 pixel high one-channel image, and the output as a 1 pixel high 200-channel image. So what we need in between is a series of convolutions and pooling operations (i.e., a fully-convolutional network) that reduces the image height from 80 to 1, and produces 200 channels.<br>\nFor illustrative purposes, assume we use a single convolutional layer for that: It would need 200 filters of height 80, but we are still free to choose the width (corresponding to the temporal context used for a single local prediction), and the horizontal stride (corresponding to the temporal density of predictions, i.e., we may find that we do not need predictions at the same rate as our 70 spectrogram frames per second).<br>\nOf course, we went deeper than a single layer, but we can still compute the receptive field width (the temporal context used for a prediction) and the horizontal stride (the prediction rate) for a fully-convolutional network.</p>\n<p>I experimented with three different architectures: A \"vanilla\" ConvNet, a small residual network, and a pretrained PANN.</p>\n<h3>Vanilla</h3>\n<p>The most simple architecture (based on \"sparrow\" from a <a href=\"http://ofai.at/~jan.schlueter/pubs/2017_eusipco.pdf\" target=\"_blank\">2017 EUSIPCO paper of ours</a>) was <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/defaults.vars#L21\" target=\"_blank\">specified in our code</a> as follows:</p>\n<pre><code>conv2d:64@3x3,bn2d,lrelu,\nconv2d:64@3x3,bn2d,\npool2d:max@3x3,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@17x3,bn2d,\npool2d:max@5x3,lrelu,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\n</code></pre>\n<p>Where</p>\n<ul>\n<li><code>conv2d:F@HxW</code> denotes an unpadded 2d convolution of <code>F</code> filters with height <code>H</code> (frequency bands) and width <code>W</code> (time frames),</li>\n<li><code>pool2d:max@HxW</code> denotes non-overlapping unpadded 2d max pooling of height <code>H</code> and width <code>W</code>,</li>\n<li><code>bn2d</code> denotes batch normalization,</li>\n<li><code>lrelu</code> denotes the leaky rectifier (LReLU) of leakiness 0.01,</li>\n<li>and <code>dropout:0.5</code> denotes ordinary (non-spatial) dropout of 50%.</li>\n</ul>\n<p>The last convolution is indicated with <code>C</code> channels, this will be replaced by 200, the number of classes. (The last two convolutions can be seen as two fully-connected layers of 1024 and 200 dimensions operating on each time step.)</p>\n<p>In total, this has a receptive field of 79x103 (79 frequency bands, 103 spectrogram frames, about 1.5 seconds) and a temporal stride of 9 (~7.778 predictions per second). (Late insight: We will actually lose the highest frequency band due to the first pooling operation – we could have used a mel filterbank of 79 filters up to 9646.4 Hz instead for this model.) It clocks in at 3.6 mio parameters.</p>\n<p>As our batch size was not that large (see the section on training), I later replaced batch normalization with group normalization, using 16 groups throughout the model. This improved performance a bit.</p>\n<h3>ResNet</h3>\n<p>As in <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">BirdCLEF 2018</a>, I extended the previous model into a simple residual network by replacing each of the first four convolutions with a residual block. Taken <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L170\" target=\"_blank\">from our config file</a>:</p>\n<pre><code>add[conv2d:64@3x3,bn2d,relu,\n    conv2d:64@3x3\n   |crop2d:2,conv2d:64@1x1],\nadd[bn2d,relu,conv2d:64@3x3,\n    bn2d,relu,conv2d:64@3x3\n   |crop2d:2],\npool2d:max@3x3,\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2,conv2d:128@1x1],\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2],\nbn2d,relu,\nconv2d:128@12x3,bn2d,\nlrelu,pool2d:max@5x3,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\"\n</code></pre>\n<p>Where:</p>\n<ul>\n<li><code>add[A|B]</code> adds the output of two branches, and</li>\n<li><code>crop2d:K</code> reduces the height and width by <code>K</code> pixels on each side, needed to match what is lost in the unpadded convolutions.</li>\n</ul>\n<p>The receptive field is at 80x119, only minimally larger, and the temporal stride stays at 9 frames. With 3.73 mio parameters it is a little bit larger than the vanilla model, it trains longer, and performs a bit better.</p>\n<p>Again, it later turned out this worked better with group normalization.</p>\n<h3>PANN</h3>\n<p>For this challenge, it was explicitly allowed to use models pretrained on external data, as long as they were disclosed <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\" target=\"_blank\">on the forum</a>. Among the models listed there is <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">Qiuqiang Kong's PANN repository</a> of CNNs trained on AudioSet, which is also used in <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">Hidehi Saarai's notebook</a>.</p>\n<p>I downloaded <a href=\"https://zenodo.org/record/3987831\" target=\"_blank\">the weights</a> for his <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2542\" target=\"_blank\">Cnn14_16k</a> model and <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L58\" target=\"_blank\">reproduced its central part</a> to stack it onto my frontend.</p>\n<p>This required some adaptations to the frontend to produce spectrograms compatible with what the PANN expects. I could mostly adopt <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2547-L2552\" target=\"_blank\">the settings from the PANN repository</a>, but as my inputs are sampled at 22050 Hz instead of 16000 Hz, I had to <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L92-L94\" target=\"_blank\">scale up the FFT window and hop size accordingly</a>. In addition, I needed to <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L108\" target=\"_blank\">scale and shift</a> the result of my magnitude compression. I verified that remaining differences in the formulation of the mel filterbank did not affect predictions on an AudioSet example.</p>\n<p>On top of Cnn14_16k`s 6 convolutional blocks (ending in 2048 feature maps; using only the first 4 or 5 blocks resulted in worse performance), I added two new 1x1 convolutions of 1024 and 200 channels each, with a leaky rectifier in between and 50% ordinary dropout applied to each.</p>\n<p>Overall, this model has a receptive field of 284x284 and a stride of 32x32, on spectrograms of 64 frequency bands and 100 frames per second (so it produces ~3 predictions per second, taking 2.84 seconds of context into account for a prediction, and its predictions still have 2 frequency bands that will be taken care of in the global pooling step). The receptive field of 284 frequency bands on a spectrogram of only 64 bands is achieved by excessive zero-padding (via padded convolutions).</p>\n<p>This model is much larger (~78 mio parameters) and slower to train, but achieves notably better results. This is in part since it is pretrained on AudioSet (training from scratch instead of using the pretrained weights reduced F1 score on the validation set from 0.688 to 0.676 in an early experiment).</p>\n<h2>Global pooling</h2>\n<p>Up to here, our model produces a time series of logits for each class. The final step is to pool these logits into a single prediction per class for the full recording, such that we can compute (and minimize) the classification error wrt. the given global labels for the recording.</p>\n<p>Reproducing the reasoning from my <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">BirdCLEF 2018 model</a>, there are two obvious ways and a third one in between:</p>\n<ul>\n<li><strong>mean pooling:</strong> This computes the average of the predictions over time, separately per class. The effect is that our prediction will be more confident for a bird that is detected many times during the recording compared to a bird that is detected once (even if very confidently).<br>\nThis is a bad idea: the label for a recording should reflect whether a bird is present, irrespective of how often it can be heard.<br>\nMean pooling will also distribute the gradient of the loss uniformly over all time points, training the network to predict each labeled species everywhere in the recording.</li>\n<li><strong>max pooling:</strong> This computes the maximum over time per class, so a single confident local detection will result in a confident global detection, matching the meaning of the global annotations.<br>\nHowever, max pooling will backpropagate the gradient of the loss only to the single time point where the prediction was most confident (separately for each species), pulling it up if the bird appears among the global labels, and pushing it down otherwise. This means a lot of computation for a very sparse update which ignores all the other vocalizations of the bird in the same recording.</li>\n<li><strong>log-mean-exp pooling:</strong> As a compromise, \\(\\frac{1}{a} \\log \\left( \\frac{1}{T} \\sum_{t=1}^{T} \\exp (a x_t) \\right) \\) allows to interpolate between taking the maximum (\\(a \\rightarrow \\infty\\)) and mean (\\(a \\rightarrow 0\\)). With \\(a=1\\), the output depends on the largest couple of values, which is also where the gradient is distributed to. This is what I used for most models. I also tried learning it, and learning a separate \\(a\\) per species (since some species might vocalize densely, warranting a small \\(a\\), and others sparsely, requiring a large \\(a\\)), but this only improved my validation scores without helping on the challenge data.</li>\n</ul>\n<p>Note that pooling operates in the logit domain here, that is, before applying the sigmoid that turns predictions into probabilities. This does not make a difference in max-pooling, but otherwise I empirically found that this improves results.</p>",
      "rawMarkdown": "Team\n----\n\nWe are three PhD students (Khaled Koutini, Lukas Martak, Paul Primus) and a postdoc (Jan Schlüter) from the [Institute of Computational Perception](https://https://www.jku.at/en/institute-of-computational-perception/) at Johannes-Kepler-University Linz, Austria. Since 20 of our 23 submissions were done by me (Jan), most of this is written from my perspective, but each of us did important contributions I will highlight below.\n\n\nNetwork architecture\n--------------------\n\nThe general model architecture looks as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2Faf6861f1136591254f287a776e3682f1%2Fwriteup_network.png?generation=1608311897835230&alt=media)\n\n\nFrom an (arbitrarily long) monophonic raw audio recording, the part denoted as \"Frontend\" computes a spectrogram-like representation. In the next step, a Fully-Convolutional Network (FCN) processes this representation into a time series of logits for every class. When passed through a sigmoid, these would give us local predictions at every time step. Since we do not have local labels to train these, only file-wise labels, we apply a global pooling operation (over time) to obtain a single logit per class. Passed through a sigmoid, these serve as our file-wise predictions.\n\nThis layout matches what we used for [BirdCLEF 2018](https://github.com/f0k/birdclef2018), except that we are using a sigmoid instead of a softmax now (in BirdCLEF, evaluation was based on the ranking of species probabilities, for which a softmax helped). It also matches [what Hidehisa Arai explained](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection) early on in the competition.\n\nWe still have several options for the three components: The frontend, the local predictor, and the global pooling operation.\n\n### Frontend\n\n<details>\n<summary>PyTorch architecture listing:</summary>\n\n```\n(frontend): Sequential(\n  (filterbank): Sequential(\n    (stft): STFT(winsize=1024, hopsize=315, complex=False)\n    (melfilter): MelFilter(num_bands=80, min_freq=27.5, max_freq=10000.0)\n  )\n  (magscale): Log1p(trainable=True)\n  (denoise): SubtractMedian()\n  (norm): TemporalBatchNorm(\n    (bn): BatchNorm1d(80, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n  )\n)\n```\n</details>\n\nThe frontend takes in audio recordings at a sample rate of 22050 Hz. (While the test recordings were at 32000 Hz, not all training recordings go that high, so I resampled everything to 22050 Hz.)\n\nIt consists of:\n- [STFT](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L171):\n  window size of 1024 samples, hop size of 315 samples (resulting in 22050/315=70 frames per second), with Hann window, keeping only the magnitudes.\n- [mel filterbank](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L268):\n  80 triangular filters from 27.5 Hz to 10 kHz (not going up to 11025 Hz on purpose to leave some room for pitch shifting)\n- [nonlinear magnitude scaling](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L72):\n  Magnitudes are compressed by passing them through \\\\(y = \\log(1 + 10^a x)\\\\), where \\\\(a\\\\) is initialized to zero and learned by backpropagation.\n  I got similar, but maybe slightly worse results with \\\\(y = x^{\\sigma(a)}\\\\), where \\\\(a\\\\) is initialized to zero (resulting in \\\\(y = \\sqrt{x}\\\\)) and learned by backpropagation.\n- [denoising](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L163):\n  The different recordings have very different background noise floors, both due to the different environments and due to different recording equipment. Human listeners are quite good at adapting to this within a few seconds. To make it easier for the network, I unify the recordings somewhat by subtracting the median over time from each frequency band (separately for each recording or excerpt). I also tried [Per-Channel Energy Normalization](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L114), but subtracting the median worked better (and is faster). For PANNs this step is skipped, as it wouldn't match the input they are pretrained on.\n- [normalization](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L53):\n  To ensure inputs are in a reasonable range (and stay in a reasonable range when the magnitude scaling changes during training), each frequency band is normalized over time with Batch Normalization.\n\n### Local predictor\n\nThe purpose of the local predictor is to take the spectrogram produced by the frontend, and produce 200 time series of logits, one for each bird species.\nThe spectrogram can be regarded as an 80 pixel high one-channel image, and the output as a 1 pixel high 200-channel image. So what we need in between is a series of convolutions and pooling operations (i.e., a fully-convolutional network) that reduces the image height from 80 to 1, and produces 200 channels.\nFor illustrative purposes, assume we use a single convolutional layer for that: It would need 200 filters of height 80, but we are still free to choose the width (corresponding to the temporal context used for a single local prediction), and the horizontal stride (corresponding to the temporal density of predictions, i.e., we may find that we do not need predictions at the same rate as our 70 spectrogram frames per second).\nOf course, we went deeper than a single layer, but we can still compute the receptive field width (the temporal context used for a prediction) and the horizontal stride (the prediction rate) for a fully-convolutional network.\n\nI experimented with three different architectures: A \"vanilla\" ConvNet, a small residual network, and a pretrained PANN.\n\n### Vanilla\n\nThe most simple architecture (based on \"sparrow\" from a [2017 EUSIPCO paper of ours](http://ofai.at/~jan.schlueter/pubs/2017_eusipco.pdf)) was [specified in our code](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/defaults.vars#L21) as follows:\n```\nconv2d:64@3x3,bn2d,lrelu,\nconv2d:64@3x3,bn2d,\npool2d:max@3x3,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@17x3,bn2d,\npool2d:max@5x3,lrelu,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\n```\nWhere\n* `conv2d:F@HxW` denotes an unpadded 2d convolution of `F` filters with height `H` (frequency bands) and width `W` (time frames),\n* `pool2d:max@HxW` denotes non-overlapping unpadded 2d max pooling of height `H` and width `W`,\n* `bn2d` denotes batch normalization,\n* `lrelu` denotes the leaky rectifier (LReLU) of leakiness 0.01,\n* and `dropout:0.5` denotes ordinary (non-spatial) dropout of 50%.\n\nThe last convolution is indicated with `C` channels, this will be replaced by 200, the number of classes. (The last two convolutions can be seen as two fully-connected layers of 1024 and 200 dimensions operating on each time step.)\n\nIn total, this has a receptive field of 79x103 (79 frequency bands, 103 spectrogram frames, about 1.5 seconds) and a temporal stride of 9 (~7.778 predictions per second). (Late insight: We will actually lose the highest frequency band due to the first pooling operation &ndash; we could have used a mel filterbank of 79 filters up to 9646.4 Hz instead for this model.) It clocks in at 3.6 mio parameters.\n\nAs our batch size was not that large (see the section on training), I later replaced batch normalization with group normalization, using 16 groups throughout the model. This improved performance a bit.\n\n### ResNet\n\nAs in [BirdCLEF 2018](https://github.com/f0k/birdclef2018), I extended the previous model into a simple residual network by replacing each of the first four convolutions with a residual block. Taken [from our config file](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L170):\n```\nadd[conv2d:64@3x3,bn2d,relu,\n    conv2d:64@3x3\n   |crop2d:2,conv2d:64@1x1],\nadd[bn2d,relu,conv2d:64@3x3,\n    bn2d,relu,conv2d:64@3x3\n   |crop2d:2],\npool2d:max@3x3,\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2,conv2d:128@1x1],\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2],\nbn2d,relu,\nconv2d:128@12x3,bn2d,\nlrelu,pool2d:max@5x3,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\"\n```\nWhere:\n* `add[A|B]` adds the output of two branches, and\n* `crop2d:K` reduces the height and width by `K` pixels on each side, needed to match what is lost in the unpadded convolutions.\n\nThe receptive field is at 80x119, only minimally larger, and the temporal stride stays at 9 frames. With 3.73 mio parameters it is a little bit larger than the vanilla model, it trains longer, and performs a bit better.\n\nAgain, it later turned out this worked better with group normalization.\n\n### PANN\n\nFor this challenge, it was explicitly allowed to use models pretrained on external data, as long as they were disclosed [on the forum](https://www.kaggle.com/c/birdsong-recognition/discussion/158877). Among the models listed there is [Qiuqiang Kong's PANN repository](https://github.com/qiuqiangkong/audioset_tagging_cnn) of CNNs trained on AudioSet, which is also used in [Hidehi Saarai's notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection).\n\nI downloaded [the weights](https://zenodo.org/record/3987831) for his [Cnn14_16k](https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2542) model and [reproduced its central part](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L58) to stack it onto my frontend.\n\nThis required some adaptations to the frontend to produce spectrograms compatible with what the PANN expects. I could mostly adopt [the settings from the PANN repository](https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2547-L2552), but as my inputs are sampled at 22050 Hz instead of 16000 Hz, I had to [scale up the FFT window and hop size accordingly](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L92-L94). In addition, I needed to [scale and shift](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L108) the result of my magnitude compression. I verified that remaining differences in the formulation of the mel filterbank did not affect predictions on an AudioSet example.\n\nOn top of Cnn14_16k`s 6 convolutional blocks (ending in 2048 feature maps; using only the first 4 or 5 blocks resulted in worse performance), I added two new 1x1 convolutions of 1024 and 200 channels each, with a leaky rectifier in between and 50% ordinary dropout applied to each.\n\nOverall, this model has a receptive field of 284x284 and a stride of 32x32, on spectrograms of 64 frequency bands and 100 frames per second (so it produces ~3 predictions per second, taking 2.84 seconds of context into account for a prediction, and its predictions still have 2 frequency bands that will be taken care of in the global pooling step). The receptive field of 284 frequency bands on a spectrogram of only 64 bands is achieved by excessive zero-padding (via padded convolutions).\n\nThis model is much larger (~78 mio parameters) and slower to train, but achieves notably better results. This is in part since it is pretrained on AudioSet (training from scratch instead of using the pretrained weights reduced F1 score on the validation set from 0.688 to 0.676 in an early experiment).\n\n## Global pooling\n\nUp to here, our model produces a time series of logits for each class. The final step is to pool these logits into a single prediction per class for the full recording, such that we can compute (and minimize) the classification error wrt. the given global labels for the recording.\n\nReproducing the reasoning from my [BirdCLEF 2018 model](https://github.com/f0k/birdclef2018), there are two obvious ways and a third one in between:\n* **mean pooling:** This computes the average of the predictions over time, separately per class. The effect is that our prediction will be more confident for a bird that is detected many times during the recording compared to a bird that is detected once (even if very confidently).\nThis is a bad idea: the label for a recording should reflect whether a bird is present, irrespective of how often it can be heard.\nMean pooling will also distribute the gradient of the loss uniformly over all time points, training the network to predict each labeled species everywhere in the recording.\n* **max pooling:** This computes the maximum over time per class, so a single confident local detection will result in a confident global detection, matching the meaning of the global annotations.\nHowever, max pooling will backpropagate the gradient of the loss only to the single time point where the prediction was most confident (separately for each species), pulling it up if the bird appears among the global labels, and pushing it down otherwise. This means a lot of computation for a very sparse update which ignores all the other vocalizations of the bird in the same recording.\n* **log-mean-exp pooling:** As a compromise, \\\\(\\frac{1}{a} \\log \\left( \\frac{1}{T} \\sum_{t=1}^{T} \\exp (a x_t) \\right) \\\\) allows to interpolate between taking the maximum (\\\\(a \\rightarrow \\infty\\\\)) and mean (\\\\(a \\rightarrow 0\\\\)). With \\\\(a=1\\\\), the output depends on the largest couple of values, which is also where the gradient is distributed to. This is what I used for most models. I also tried learning it, and learning a separate \\\\(a\\\\) per species (since some species might vocalize densely, warranting a small \\\\(a\\\\), and others sparsely, requiring a large \\\\(a\\\\)), but this only improved my validation scores without helping on the challenge data.\n\nNote that pooling operates in the logit domain here, that is, before applying the sigmoid that turns predictions into probabilities. This does not make a difference in max-pooling, but otherwise I empirically found that this improves results.",
      "votes": null
    },
    {
      "id": "1121413",
      "postDate": "12/21/2020 16:18:58",
      "content": "<h2>Data</h2>\n<p>We used both the official training data and <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">Rohan Rao's xeno-canto crawls</a>, totalling in 43499 recordings for the 200 species. We reserved about 10% for validation (splitting such that no recordist is part of both the training and validation set). Labels contain both the annotated foreground and background species for a recording, without distinguishing them (setting lower target values or lower weights for background species did not improve results).</p>\n<p>All recordings are predecoded to 22050 Hz 16-bit stereo or mono wave files. As part of the data loading pipeline, stereo files are <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L259\" target=\"_blank\">downmixed to mono</a> with a randomly uniform weight \\(p\\) for the left channel, and \\(1-p\\) for the right channel. This augmentation slightly improves results over downmixing them with fixed weights during decoding.</p>\n<p>To make the models work under low signal-to-noise ratios (i.e., the conditions found in the test files), we mix them with excerpts from the <a href=\"https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591\" target=\"_blank\">Chernobyl BiVA</a> and <a href=\"https://zenodo.org/record/1205569\" target=\"_blank\">BirdVox-full-night</a> datasets, which Paul suggested and hunted down. These datasets are precisely annotated with bird occurrences, so we can extract all parts void of birds. Mixing is <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L195\" target=\"_blank\">done on-the-fly</a> as part of the data loader. I started out carefully, but the best setting turned out to be mixing <em>every</em> training example with background noise, drawing a value \\(p \\in [0,1)\\) and scaling the noise with \\(p\\) and the bird recording excerpt with \\(1-p\\). For two of the models in the final ensemble, I also set \\(p=1\\) with 1% probability, setting the labels to all zero in this case.</p>\n<p>For model selection (explained further below), we also made use of the two official <code>example_test_audio</code> recordings, as well as the six 10-minute North American BirdCLEF 2020 validation set recordings <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911091\" target=\"_blank\">posted on the forum</a>. However, they only contain a fraction of the relevant species, contain species not part of the competition, and are not accurately labeled (missing several calls).</p>\n<h2>Training</h2>\n<p>Ideally, the model would be trained on complete recordings – this is the only way we can be sure all the targets are correct. If we pick a random excerpt, it is not guaranteed that all birds annotated to be present in the recording are also audible in the chosen excerpt. However, the longest recordings are 3 hours, which is impractical. As for BirdCLEF 2018, we train on randomly selected 30-second excerpts instead, hoping that most annotated birds will be audible at least once. Too short files are looped to make up 30 seconds (for BirdCLEF, we selected mini-batches among files of similar length to avoid unneeded looping, but PyTorch runs out of memory if the input tensors change shapes all the time). Validation uses the central 30 seconds of a recording.</p>\n<p>Training uses ADAM with mini-batches of 16 examples, an initial learning rate of 1e-3, and PyTorch's default settings for beta1, beta2 and epsilon. The validation loss is computed every 1000 update steps. If it does not improve over the current best value for 10 such evaluations in a row, the learning rate is reduced to a tenth, and training is continued. Training is stopped when the learning rate reaches 1e-6.</p>\n<p>For the PANN, I tried reducing the learning rate for the pretrained layers to 1% or 10% compared to the novel layers or to freeze the pretrained layers for some time, but it turned out that using the full learning rate for all the layers from the start works best.</p>\n<p>Training takes about 6h for a vanilla model, 7h for a residual network, and 8h for a PANN, on an RTX 2080 Ti.</p>\n<h2>Inference</h2>\n<p>The challenge required two types of inference:</p>\n<ol>\n<li>Predicting a set of species for every non-overlapping 5-second window of a 10-minute recording (for recording sites 1 and 2)</li>\n<li>Predicting a set of species for a full 10-minute recording (for recording site 3)</li>\n</ol>\n<p>Since the network is built to produce local predictions that are then pooled over time, we can simply pass a complete recording through the network up to the pooling operation, and then either</p>\n<ol>\n<li>Apply pooling with non-overlapping 5-second windows</li>\n<li>Apply pooling over the full recording</li>\n</ol>\n<p>The first case requires some trickery, since the network produces an uneven number of predictions within 5 seconds (5x70/9=38.889 for the vanilla model and ResNet, and 5x100/32=15.625 for the PANN). I solved it by having the models compute and report their receptive field (size, padding and stride) wrt. the audio samples, using this to <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L252-L258\" target=\"_blank\">compute a time point for each prediction</a>, and then <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L87\" target=\"_blank\">placing the pooling boundaries</a> appropriately. Note that this produces similar results as presegmenting the audio into 5-second windows and passing them through the network, but without artifacts introduced by zero-padding for the PANN, and with considerably less computation.</p>\n<p>Unfortunately, both types of inference do not match what the model was trained for: to produce good labels for 30-second audio excerpts. Since we use global log-mean-exp pooling, the output of the pooling stage is not completely independent of the input length; it is still somewhat resembling a global mean pooling:</p>\n<ol>\n<li>A short false positive that would stay under the threshold in a 30-second excerpt may score above the threshold for a 5-second window.</li>\n<li>If we feed 10-minute recordings, birds that are locally detected in some parts of the recording may still fall under the threshold after pooling.</li>\n</ol>\n<p>After 8 days of optimizing the models (4.5 days before the end of the challenge), I changed inference for the second case:</p>\n<p>&nbsp; &nbsp; 2. Apply pooling with 20-second windows overlapping by 50%, then take the maximum over the pooled windows for each species.</p>\n<p>Using 20 seconds instead of 30 seconds seemed to produce slightly better results on the validation data. However, this new scheme did not change my score on the public leaderboard – probably because it only affects the recordings from site 3, which may have a small influence on the total score (depending on how they are weighted). (On the private set, the score improved from 0.656 to 0.657, also unimportant.)</p>\n<p>After some more experimentation with models (improving the score to 0.587 by switching to PANNs), one day and 15 minutes before the deadline, I also changed inference for the first case (recording sites 1 and 2):</p>\n<ol>\n<li>First predict species globally for the 10-minute recording, using the new inference scheme developed for site 3. Then predict species locally by pooling in non-overlapping 5-second windows, but <strong>using a detection threshold of 0.3 instead of 0.5</strong>, and <strong>only keeping species that were predicted globally</strong> for that recording.</li>\n</ol>\n<p>The idea is that we should use longer windows to reliably detect birds, but once we are sure they are present in a recording, we can increase the classifier's sensitivity to find all occurrences.<br>\nThis gave a real boost from 0.587 to 0.600 on the public leaderboard. I thought about spending the second submission of that day on reducing the threshold to 0.2, but tried submitting an improved ensemble instead, which then finished uploading 30 seconds too late.</p>\n<p>Regarding computational efficiency, the final ensemble consisting of 3 PANNs took 4 minutes to run over the test set on Kaggle (but I do not know if and by how much the submission notebook evaluation is parallellized over subsets of the data).</p>\n<h2>Model selection</h2>\n<p>Most of our submissions used ensembles of 3 to 5 models selected by their F1 score on the validation set (10% of the training data from xeno-canto). The validation set may not be a good indicator for the challenge test set performance since a) it consists of focal recordings, not soundscapes, and b) it only comes with file-wise labels, not 5-second chunks or smaller. In addition, c) its distribution of species probably differs significantly from the test data.</p>\n<p>In addition to the validation set, we used the 8 annotated soundscape excerpts (2 from the challenge dataset, 6 from BirdCLEF 2020) to verify aspects such as the inference strategy. The 8 soundscape excerpts are also not a good indicator for the challenge performance since a) they only cover a small subset of the species and b) they are not thoroughly annotated.</p>\n<p>On the final day, I tried to improve model selection by setting up 6 different tasks:</p>\n<ol>\n<li>Prediction in 5-second windows on the 8 annotated soundscape excerpts</li>\n<li>Predicting file-level labels for the 8 annotated soundscape excerpts</li>\n<li>Predicting file-level labels on the validation set</li>\n<li>As 3., but recordings augmented with background noise from Chernobyl and BirdVox</li>\n<li>Predicting file-level labels only on those files from the validation set which have annotated background species (the other files may not be annotated completely)</li>\n<li>As 5., but recordings augmented with background noise from Chernobyl and BirdVox</li>\n</ol>\n<p>I evaluated all 16 previous submissions on these 6 tasks and looked for some correlation of precision, recall or F1 score with the submissions' public leaderboard performance:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F50bc581848f305ec9036bbca2313efed%2Fwriteup_correlation.png?generation=1608567691969180&amp;alt=media\"><br>\nThere was no clear indicator. With a combination of two scores that was at least almost monotonically increasing (when ordered by public leaderboard scores), I selected an ensemble that improved public leaderboard score from 0.600 to 0.601.</p>\n<p>Disappointed, for the last submission I chose another ensemble based on validation set performance alone. I hesitated whether to try reducing the prediction threshold from 0.3 to 0.2, but was afraid to ruin the submission. It scored 0.596 on the public leaderboard (it would have made the 4th place on the private set).</p>\n<h2>What did not work</h2>\n<ul>\n<li>Khaled ran several experiments using his <a href=\"https://arxiv.org/abs/1909.02859\" target=\"_blank\">receptive-field-regularized CNNs</a> which he successfully employed in past DCASE challenges, but unfortunately was not able to surpass the other models.</li>\n<li>Lukas contributed an <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L430\" target=\"_blank\">augmentation using colored noise</a>, but it turned out not to improve results over augmenting with the Chernobyl and BirdVox recordings. In addition, he worked on pretraining a model with self-supervised learning.</li>\n<li>Using shorter excerpts (and larger batches) made results worse, longer excerpts (and shorter batches) did not improve results.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L405-L414\" target=\"_blank\">Reducing the frequency range</a> in order to have higher frequency resolution did not improve results.</li>\n<li>Augmentation via a cheap pitch shift implemented by <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L297-L303\" target=\"_blank\">randomly warping the mel filterbank</a> did not improve results.</li>\n<li>Training <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L272-L278\" target=\"_blank\">only on high-quality recordings</a>, or <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L69-L74\" target=\"_blank\">weighting the loss by the quality rating</a> of the recording made results worse.</li>\n<li>Mixup made results worse (vanilla model, validation set results: 0.661 without mixup, 0.653 with alpha=0.1, 0.643 with alpha=0.3).</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/metrics/__init__.py#L171-L174\" target=\"_blank\">Label smoothing</a> did not have a consistent effect (PANN, validation set results: 0.690 without, 0.682 with 1% smoothing, 0.687 with 5% smoothing).</li>\n<li><a href=\"http://proceedings.mlr.press/v80/furlanello18a.html\" target=\"_blank\">Born-Again Neural Networks</a> helped for BirdCLEF 2018, but not here, at least not for the PANN (validation set results: 0.690 originally, 0.685 as a born-again network).</li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>Apart from the challenge hosts, we are indebted to the following people:</p>\n<ul>\n<li>Hidehisa Arai for his <a href=\"https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline/#Prediction-loop\" target=\"_blank\">template on computing predictions</a> that falls back to a replacement test dataset at development time, when the actual test data is not available</li>\n<li>Alex Shonenkov for compiling this <a href=\"https://www.kaggle.com/shonenkov/birdcall-check\" target=\"_blank\">replacement test dataset</a></li>\n<li>Rohan Rao for crawling xeno-canto to compile his <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">additional training dataset</a></li>\n<li>All competitors who raised the bar on the leaderboard, and everyone discussing on the forum!</li>\n</ul>\n<h2>Afterthoughts</h2>\n<p>After the challenge finished, I was anxious to check if reducing the threshold from 0.3 to 0.2 would have made a difference. Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions 😃<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F70b5a32bbc3cd1b741c6696251c54c25%2Fmoneyknob.png?generation=1600330942624036&amp;alt=media\" alt=\"\"><br>\nA threshold of 0.2 would have ended on the third place, a threshold of 0.1 or less (even 0.02) would have made the second place.</p>\n<p><em>Were my submissions close to being in the money?</em><br>\nArguably yes, the models were fine, it only took changing a single hyperparameter in the inference pipeline that I was fully aware of. But as CPMP mentioned, that's probably true of other submissions as well!</p>\n<p><em>Was I close to submitting a winning solution?</em><br>\nThis thought has given me some unrest after the competition, but I think I wasn't.<br>\nAfter my first try of the new inference strategy, there were only three submissions remaining. I thought about trying a reduced threshold, but reducing the threshold in the old inference strategy had yielded worse results (despite an improved F-score on the 8 annotated soundscapes available for testing), so I was pessimistic about this. If I had tried anyway, I probably wouldn't have spent the last two submissions on reducing it far enough to be in the money with that ensemble.<br>\nOn the last submission, minutes before the deadline, I considered submitting the last new ensemble with threshold 0.2, but was afraid to change two things at the same time and ruin it. It would have yielded 0.599 on the public set. I would have thought I ruined it and not have selected it for the final standing.<br>\nIf I had started the whole competition earlier, I might have fallen for optimizing too much on the public set and missing the top ten. So overall I'm glad how it played out!</p>",
      "rawMarkdown": "Data\n----\n\nWe used both the official training data and [Rohan Rao's xeno-canto crawls](https://www.kaggle.com/c/birdsong-recognition/discussion/159970), totalling in 43499 recordings for the 200 species. We reserved about 10% for validation (splitting such that no recordist is part of both the training and validation set). Labels contain both the annotated foreground and background species for a recording, without distinguishing them (setting lower target values or lower weights for background species did not improve results).\n\nAll recordings are predecoded to 22050 Hz 16-bit stereo or mono wave files. As part of the data loading pipeline, stereo files are [downmixed to mono](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L259) with a randomly uniform weight \\\\(p\\\\) for the left channel, and \\\\(1-p\\\\) for the right channel. This augmentation slightly improves results over downmixing them with fixed weights during decoding.\n\nTo make the models work under low signal-to-noise ratios (i.e., the conditions found in the test files), we mix them with excerpts from the [Chernobyl BiVA](https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591) and [BirdVox-full-night](https://zenodo.org/record/1205569) datasets, which Paul suggested and hunted down. These datasets are precisely annotated with bird occurrences, so we can extract all parts void of birds. Mixing is [done on-the-fly](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L195) as part of the data loader. I started out carefully, but the best setting turned out to be mixing *every* training example with background noise, drawing a value \\\\(p \\in [0,1)\\\\) and scaling the noise with \\\\(p\\\\) and the bird recording excerpt with \\\\(1-p\\\\). For two of the models in the final ensemble, I also set \\\\(p=1\\\\) with 1% probability, setting the labels to all zero in this case.\n\nFor model selection (explained further below), we also made use of the two official `example_test_audio` recordings, as well as the six 10-minute North American BirdCLEF 2020 validation set recordings [posted on the forum](https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911091). However, they only contain a fraction of the relevant species, contain species not part of the competition, and are not accurately labeled (missing several calls).\n\n\nTraining\n--------\n\nIdeally, the model would be trained on complete recordings &ndash; this is the only way we can be sure all the targets are correct. If we pick a random excerpt, it is not guaranteed that all birds annotated to be present in the recording are also audible in the chosen excerpt. However, the longest recordings are 3 hours, which is impractical. As for BirdCLEF 2018, we train on randomly selected 30-second excerpts instead, hoping that most annotated birds will be audible at least once. Too short files are looped to make up 30 seconds (for BirdCLEF, we selected mini-batches among files of similar length to avoid unneeded looping, but PyTorch runs out of memory if the input tensors change shapes all the time). Validation uses the central 30 seconds of a recording.\n\nTraining uses ADAM with mini-batches of 16 examples, an initial learning rate of 1e-3, and PyTorch's default settings for beta1, beta2 and epsilon. The validation loss is computed every 1000 update steps. If it does not improve over the current best value for 10 such evaluations in a row, the learning rate is reduced to a tenth, and training is continued. Training is stopped when the learning rate reaches 1e-6.\n\nFor the PANN, I tried reducing the learning rate for the pretrained layers to 1% or 10% compared to the novel layers or to freeze the pretrained layers for some time, but it turned out that using the full learning rate for all the layers from the start works best.\n\nTraining takes about 6h for a vanilla model, 7h for a residual network, and 8h for a PANN, on an RTX 2080 Ti.\n\n\nInference\n---------\n\nThe challenge required two types of inference:\n1. Predicting a set of species for every non-overlapping 5-second window of a 10-minute recording (for recording sites 1 and 2)\n2. Predicting a set of species for a full 10-minute recording (for recording site 3)\n\nSince the network is built to produce local predictions that are then pooled over time, we can simply pass a complete recording through the network up to the pooling operation, and then either\n1. Apply pooling with non-overlapping 5-second windows\n2. Apply pooling over the full recording\n\nThe first case requires some trickery, since the network produces an uneven number of predictions within 5 seconds (5x70/9=38.889 for the vanilla model and ResNet, and 5x100/32=15.625 for the PANN). I solved it by having the models compute and report their receptive field (size, padding and stride) wrt. the audio samples, using this to [compute a time point for each prediction](https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L252-L258), and then [placing the pooling boundaries](https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L87) appropriately. Note that this produces similar results as presegmenting the audio into 5-second windows and passing them through the network, but without artifacts introduced by zero-padding for the PANN, and with considerably less computation.\n\nUnfortunately, both types of inference do not match what the model was trained for: to produce good labels for 30-second audio excerpts. Since we use global log-mean-exp pooling, the output of the pooling stage is not completely independent of the input length; it is still somewhat resembling a global mean pooling:\n1. A short false positive that would stay under the threshold in a 30-second excerpt may score above the threshold for a 5-second window.\n2. If we feed 10-minute recordings, birds that are locally detected in some parts of the recording may still fall under the threshold after pooling.\n\nAfter 8 days of optimizing the models (4.5 days before the end of the challenge), I changed inference for the second case:\n\n&nbsp; &nbsp; 2. Apply pooling with 20-second windows overlapping by 50%, then take the maximum over the pooled windows for each species.\n\nUsing 20 seconds instead of 30 seconds seemed to produce slightly better results on the validation data. However, this new scheme did not change my score on the public leaderboard &ndash; probably because it only affects the recordings from site 3, which may have a small influence on the total score (depending on how they are weighted). (On the private set, the score improved from 0.656 to 0.657, also unimportant.)\n\nAfter some more experimentation with models (improving the score to 0.587 by switching to PANNs), one day and 15 minutes before the deadline, I also changed inference for the first case (recording sites 1 and 2):\n1. First predict species globally for the 10-minute recording, using the new inference scheme developed for site 3. Then predict species locally by pooling in non-overlapping 5-second windows, but **using a detection threshold of 0.3 instead of 0.5**, and **only keeping species that were predicted globally** for that recording.\n\nThe idea is that we should use longer windows to reliably detect birds, but once we are sure they are present in a recording, we can increase the classifier's sensitivity to find all occurrences.\nThis gave a real boost from 0.587 to 0.600 on the public leaderboard. I thought about spending the second submission of that day on reducing the threshold to 0.2, but tried submitting an improved ensemble instead, which then finished uploading 30 seconds too late.\n\nRegarding computational efficiency, the final ensemble consisting of 3 PANNs took 4 minutes to run over the test set on Kaggle (but I do not know if and by how much the submission notebook evaluation is parallellized over subsets of the data).\n\n\nModel selection\n---------------\n\nMost of our submissions used ensembles of 3 to 5 models selected by their F1 score on the validation set (10% of the training data from xeno-canto). The validation set may not be a good indicator for the challenge test set performance since a) it consists of focal recordings, not soundscapes, and b) it only comes with file-wise labels, not 5-second chunks or smaller. In addition, c) its distribution of species probably differs significantly from the test data.\n\nIn addition to the validation set, we used the 8 annotated soundscape excerpts (2 from the challenge dataset, 6 from BirdCLEF 2020) to verify aspects such as the inference strategy. The 8 soundscape excerpts are also not a good indicator for the challenge performance since a) they only cover a small subset of the species and b) they are not thoroughly annotated.\n\nOn the final day, I tried to improve model selection by setting up 6 different tasks:\n1. Prediction in 5-second windows on the 8 annotated soundscape excerpts\n2. Predicting file-level labels for the 8 annotated soundscape excerpts\n3. Predicting file-level labels on the validation set\n4. As 3., but recordings augmented with background noise from Chernobyl and BirdVox\n5. Predicting file-level labels only on those files from the validation set which have annotated background species (the other files may not be annotated completely)\n6. As 5., but recordings augmented with background noise from Chernobyl and BirdVox\n\nI evaluated all 16 previous submissions on these 6 tasks and looked for some correlation of precision, recall or F1 score with the submissions' public leaderboard performance:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F50bc581848f305ec9036bbca2313efed%2Fwriteup_correlation.png?generation=1608567691969180&alt=media\" width=\"300\">\nThere was no clear indicator. With a combination of two scores that was at least almost monotonically increasing (when ordered by public leaderboard scores), I selected an ensemble that improved public leaderboard score from 0.600 to 0.601.\n\nDisappointed, for the last submission I chose another ensemble based on validation set performance alone. I hesitated whether to try reducing the prediction threshold from 0.3 to 0.2, but was afraid to ruin the submission. It scored 0.596 on the public leaderboard (it would have made the 4th place on the private set).\n\n\nWhat did not work\n-----------------\n\n* Khaled ran several experiments using his [receptive-field-regularized CNNs](https://arxiv.org/abs/1909.02859) which he successfully employed in past DCASE challenges, but unfortunately was not able to surpass the other models.\n* Lukas contributed an [augmentation using colored noise](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L430), but it turned out not to improve results over augmenting with the Chernobyl and BirdVox recordings. In addition, he worked on pretraining a model with self-supervised learning.\n* Using shorter excerpts (and larger batches) made results worse, longer excerpts (and shorter batches) did not improve results.\n* [Reducing the frequency range](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L405-L414) in order to have higher frequency resolution did not improve results.\n* Augmentation via a cheap pitch shift implemented by [randomly warping the mel filterbank](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L297-L303) did not improve results.\n* Training [only on high-quality recordings](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L272-L278), or [weighting the loss by the quality rating](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L69-L74) of the recording made results worse.\n* Mixup made results worse (vanilla model, validation set results: 0.661 without mixup, 0.653 with alpha=0.1, 0.643 with alpha=0.3).\n* [Label smoothing](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/metrics/__init__.py#L171-L174) did not have a consistent effect (PANN, validation set results: 0.690 without, 0.682 with 1% smoothing, 0.687 with 5% smoothing).\n* [Born-Again Neural Networks](http://proceedings.mlr.press/v80/furlanello18a.html) helped for BirdCLEF 2018, but not here, at least not for the PANN (validation set results: 0.690 originally, 0.685 as a born-again network).\n\n\nAcknowledgements\n----------------\n\nApart from the challenge hosts, we are indebted to the following people:\n* Hidehisa Arai for his [template on computing predictions](https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline/#Prediction-loop) that falls back to a replacement test dataset at development time, when the actual test data is not available\n* Alex Shonenkov for compiling this [replacement test dataset](https://www.kaggle.com/shonenkov/birdcall-check)\n* Rohan Rao for crawling xeno-canto to compile his [additional training dataset](https://www.kaggle.com/c/birdsong-recognition/discussion/159970)\n* All competitors who raised the bar on the leaderboard, and everyone discussing on the forum!\n\n\nAfterthoughts\n-------------\n\nAfter the challenge finished, I was anxious to check if reducing the threshold from 0.3 to 0.2 would have made a difference. Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions 😃\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F70b5a32bbc3cd1b741c6696251c54c25%2Fmoneyknob.png?generation=1600330942624036&alt=media)\nA threshold of 0.2 would have ended on the third place, a threshold of 0.1 or less (even 0.02) would have made the second place.\n\n*Were my submissions close to being in the money?*\nArguably yes, the models were fine, it only took changing a single hyperparameter in the inference pipeline that I was fully aware of. But as CPMP mentioned, that's probably true of other submissions as well!\n\n*Was I close to submitting a winning solution?*\nThis thought has given me some unrest after the competition, but I think I wasn't.\nAfter my first try of the new inference strategy, there were only three submissions remaining. I thought about trying a reduced threshold, but reducing the threshold in the old inference strategy had yielded worse results (despite an improved F-score on the 8 annotated soundscapes available for testing), so I was pessimistic about this. If I had tried anyway, I probably wouldn't have spent the last two submissions on reducing it far enough to be in the money with that ensemble.\nOn the last submission, minutes before the deadline, I considered submitting the last new ensemble with threshold 0.2, but was afraid to change two things at the same time and ruin it. It would have yielded 0.599 on the public set. I would have thought I ruined it and not have selected it for the final standing.\nIf I had started the whole competition earlier, I might have fallen for optimizing too much on the public set and missing the top ten. So overall I'm glad how it played out!",
      "votes": null
    },
    {
      "id": "1121631",
      "postDate": "12/21/2020 19:25:16",
      "content": "<p>Thanks for the thorough writeup! </p>\n<p>I'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.</p>\n<p>(Also, the <a href=\"url\" target=\"_blank\">http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf</a> is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)</p>",
      "rawMarkdown": "Thanks for the thorough writeup! \n\nI'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.\n\n(Also, the [http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf](url) is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)",
      "votes": null
    },
    {
      "id": "1144820",
      "postDate": "01/08/2021 17:32:37",
      "content": "<blockquote>\n  <p>I'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.</p>\n</blockquote>\n<p>I think I tried that some time and it produced worse results, but that was on a different dataset. Maybe I should try again, I'm curious now. I'll post here if I do.</p>\n<blockquote>\n  <p>(Also, the <a href=\"http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf\" target=\"_blank\">http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf</a> is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)</p>\n</blockquote>\n<p>Yes, I'm aware of it, but it's a good pointer! I only learned about it after BirdCLEF 2018. Their softmax-based pooling is highly related to log-mean-exp pooling, the latter was just around longer (Equation 6 in <a href=\"http://openaccess.thecvf.com/content_cvpr_2015/html/Pinheiro_From_Image-Level_to_2015_CVPR_paper.html\" target=\"_blank\">Pinheiro et al., 2015</a>).</p>",
      "rawMarkdown": "> I'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.\n\nI think I tried that some time and it produced worse results, but that was on a different dataset. Maybe I should try again, I'm curious now. I'll post here if I do.\n\n> (Also, the http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)\n\nYes, I'm aware of it, but it's a good pointer! I only learned about it after BirdCLEF 2018. Their softmax-based pooling is highly related to log-mean-exp pooling, the latter was just around longer (Equation 6 in [Pinheiro et al., 2015](http://openaccess.thecvf.com/content_cvpr_2015/html/Pinheiro_From_Image-Level_to_2015_CVPR_paper.html)).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1014359,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/17/2020 11:31:27",
      "content": "<blockquote>\n  <p>Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions</p>\n</blockquote>\n<p>Unfortunately this is true for many teams in every competition.  Looking for your writeup, and congrats on the result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1014384,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "09/17/2020 11:59:22",
      "content": "<p>Congratulations 🎊 </p>\n<p>Hope for next time…. Waiting for your writeup </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1121403,
      "author_name": "janschl",
      "author_url": "",
      "post_date": "12/21/2020 16:08:54",
      "content": "<h2>Team</h2>\n<p>We are three PhD students (Khaled Koutini, Lukas Martak, Paul Primus) and a postdoc (Jan Schlüter) from the <a href=\"https://https://www.jku.at/en/institute-of-computational-perception/\" target=\"_blank\">Institute of Computational Perception</a> at Johannes-Kepler-University Linz, Austria. Since 20 of our 23 submissions were done by me (Jan), most of this is written from my perspective, but each of us did important contributions I will highlight below.</p>\n<h2>Network architecture</h2>\n<p>The general model architecture looks as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2Faf6861f1136591254f287a776e3682f1%2Fwriteup_network.png?generation=1608311897835230&amp;alt=media\" alt=\"\"></p>\n<p>From an (arbitrarily long) monophonic raw audio recording, the part denoted as \"Frontend\" computes a spectrogram-like representation. In the next step, a Fully-Convolutional Network (FCN) processes this representation into a time series of logits for every class. When passed through a sigmoid, these would give us local predictions at every time step. Since we do not have local labels to train these, only file-wise labels, we apply a global pooling operation (over time) to obtain a single logit per class. Passed through a sigmoid, these serve as our file-wise predictions.</p>\n<p>This layout matches what we used for <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">BirdCLEF 2018</a>, except that we are using a sigmoid instead of a softmax now (in BirdCLEF, evaluation was based on the ranking of species probabilities, for which a softmax helped). It also matches <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">what Hidehisa Arai explained</a> early on in the competition.</p>\n<p>We still have several options for the three components: The frontend, the local predictor, and the global pooling operation.</p>\n<h3>Frontend</h3>\n<p><br>\nPyTorch architecture listing:</p>\n<pre><code>(frontend): Sequential(\n  (filterbank): Sequential(\n    (stft): STFT(winsize=1024, hopsize=315, complex=False)\n    (melfilter): MelFilter(num_bands=80, min_freq=27.5, max_freq=10000.0)\n  )\n  (magscale): Log1p(trainable=True)\n  (denoise): SubtractMedian()\n  (norm): TemporalBatchNorm(\n    (bn): BatchNorm1d(80, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n  )\n)\n</code></pre>\n<p></p>\n<p>The frontend takes in audio recordings at a sample rate of 22050 Hz. (While the test recordings were at 32000 Hz, not all training recordings go that high, so I resampled everything to 22050 Hz.)</p>\n<p>It consists of:</p>\n<ul>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L171\" target=\"_blank\">STFT</a>:<br>\nwindow size of 1024 samples, hop size of 315 samples (resulting in 22050/315=70 frames per second), with Hann window, keeping only the magnitudes.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L268\" target=\"_blank\">mel filterbank</a>:<br>\n80 triangular filters from 27.5 Hz to 10 kHz (not going up to 11025 Hz on purpose to leave some room for pitch shifting)</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L72\" target=\"_blank\">nonlinear magnitude scaling</a>:<br>\nMagnitudes are compressed by passing them through \\(y = \\log(1 + 10^a x)\\), where \\(a\\) is initialized to zero and learned by backpropagation.<br>\nI got similar, but maybe slightly worse results with \\(y = x^{\\sigma(a)}\\), where \\(a\\) is initialized to zero (resulting in \\(y = \\sqrt{x}\\)) and learned by backpropagation.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L163\" target=\"_blank\">denoising</a>:<br>\nThe different recordings have very different background noise floors, both due to the different environments and due to different recording equipment. Human listeners are quite good at adapting to this within a few seconds. To make it easier for the network, I unify the recordings somewhat by subtracting the median over time from each frequency band (separately for each recording or excerpt). I also tried <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L114\" target=\"_blank\">Per-Channel Energy Normalization</a>, but subtracting the median worked better (and is faster). For PANNs this step is skipped, as it wouldn't match the input they are pretrained on.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L53\" target=\"_blank\">normalization</a>:<br>\nTo ensure inputs are in a reasonable range (and stay in a reasonable range when the magnitude scaling changes during training), each frequency band is normalized over time with Batch Normalization.</li>\n</ul>\n<h3>Local predictor</h3>\n<p>The purpose of the local predictor is to take the spectrogram produced by the frontend, and produce 200 time series of logits, one for each bird species.<br>\nThe spectrogram can be regarded as an 80 pixel high one-channel image, and the output as a 1 pixel high 200-channel image. So what we need in between is a series of convolutions and pooling operations (i.e., a fully-convolutional network) that reduces the image height from 80 to 1, and produces 200 channels.<br>\nFor illustrative purposes, assume we use a single convolutional layer for that: It would need 200 filters of height 80, but we are still free to choose the width (corresponding to the temporal context used for a single local prediction), and the horizontal stride (corresponding to the temporal density of predictions, i.e., we may find that we do not need predictions at the same rate as our 70 spectrogram frames per second).<br>\nOf course, we went deeper than a single layer, but we can still compute the receptive field width (the temporal context used for a prediction) and the horizontal stride (the prediction rate) for a fully-convolutional network.</p>\n<p>I experimented with three different architectures: A \"vanilla\" ConvNet, a small residual network, and a pretrained PANN.</p>\n<h3>Vanilla</h3>\n<p>The most simple architecture (based on \"sparrow\" from a <a href=\"http://ofai.at/~jan.schlueter/pubs/2017_eusipco.pdf\" target=\"_blank\">2017 EUSIPCO paper of ours</a>) was <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/defaults.vars#L21\" target=\"_blank\">specified in our code</a> as follows:</p>\n<pre><code>conv2d:64@3x3,bn2d,lrelu,\nconv2d:64@3x3,bn2d,\npool2d:max@3x3,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@17x3,bn2d,\npool2d:max@5x3,lrelu,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\n</code></pre>\n<p>Where</p>\n<ul>\n<li><code>conv2d:F@HxW</code> denotes an unpadded 2d convolution of <code>F</code> filters with height <code>H</code> (frequency bands) and width <code>W</code> (time frames),</li>\n<li><code>pool2d:max@HxW</code> denotes non-overlapping unpadded 2d max pooling of height <code>H</code> and width <code>W</code>,</li>\n<li><code>bn2d</code> denotes batch normalization,</li>\n<li><code>lrelu</code> denotes the leaky rectifier (LReLU) of leakiness 0.01,</li>\n<li>and <code>dropout:0.5</code> denotes ordinary (non-spatial) dropout of 50%.</li>\n</ul>\n<p>The last convolution is indicated with <code>C</code> channels, this will be replaced by 200, the number of classes. (The last two convolutions can be seen as two fully-connected layers of 1024 and 200 dimensions operating on each time step.)</p>\n<p>In total, this has a receptive field of 79x103 (79 frequency bands, 103 spectrogram frames, about 1.5 seconds) and a temporal stride of 9 (~7.778 predictions per second). (Late insight: We will actually lose the highest frequency band due to the first pooling operation – we could have used a mel filterbank of 79 filters up to 9646.4 Hz instead for this model.) It clocks in at 3.6 mio parameters.</p>\n<p>As our batch size was not that large (see the section on training), I later replaced batch normalization with group normalization, using 16 groups throughout the model. This improved performance a bit.</p>\n<h3>ResNet</h3>\n<p>As in <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">BirdCLEF 2018</a>, I extended the previous model into a simple residual network by replacing each of the first four convolutions with a residual block. Taken <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L170\" target=\"_blank\">from our config file</a>:</p>\n<pre><code>add[conv2d:64@3x3,bn2d,relu,\n    conv2d:64@3x3\n   |crop2d:2,conv2d:64@1x1],\nadd[bn2d,relu,conv2d:64@3x3,\n    bn2d,relu,conv2d:64@3x3\n   |crop2d:2],\npool2d:max@3x3,\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2,conv2d:128@1x1],\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2],\nbn2d,relu,\nconv2d:128@12x3,bn2d,\nlrelu,pool2d:max@5x3,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\"\n</code></pre>\n<p>Where:</p>\n<ul>\n<li><code>add[A|B]</code> adds the output of two branches, and</li>\n<li><code>crop2d:K</code> reduces the height and width by <code>K</code> pixels on each side, needed to match what is lost in the unpadded convolutions.</li>\n</ul>\n<p>The receptive field is at 80x119, only minimally larger, and the temporal stride stays at 9 frames. With 3.73 mio parameters it is a little bit larger than the vanilla model, it trains longer, and performs a bit better.</p>\n<p>Again, it later turned out this worked better with group normalization.</p>\n<h3>PANN</h3>\n<p>For this challenge, it was explicitly allowed to use models pretrained on external data, as long as they were disclosed <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\" target=\"_blank\">on the forum</a>. Among the models listed there is <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">Qiuqiang Kong's PANN repository</a> of CNNs trained on AudioSet, which is also used in <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">Hidehi Saarai's notebook</a>.</p>\n<p>I downloaded <a href=\"https://zenodo.org/record/3987831\" target=\"_blank\">the weights</a> for his <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2542\" target=\"_blank\">Cnn14_16k</a> model and <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L58\" target=\"_blank\">reproduced its central part</a> to stack it onto my frontend.</p>\n<p>This required some adaptations to the frontend to produce spectrograms compatible with what the PANN expects. I could mostly adopt <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2547-L2552\" target=\"_blank\">the settings from the PANN repository</a>, but as my inputs are sampled at 22050 Hz instead of 16000 Hz, I had to <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L92-L94\" target=\"_blank\">scale up the FFT window and hop size accordingly</a>. In addition, I needed to <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L108\" target=\"_blank\">scale and shift</a> the result of my magnitude compression. I verified that remaining differences in the formulation of the mel filterbank did not affect predictions on an AudioSet example.</p>\n<p>On top of Cnn14_16k`s 6 convolutional blocks (ending in 2048 feature maps; using only the first 4 or 5 blocks resulted in worse performance), I added two new 1x1 convolutions of 1024 and 200 channels each, with a leaky rectifier in between and 50% ordinary dropout applied to each.</p>\n<p>Overall, this model has a receptive field of 284x284 and a stride of 32x32, on spectrograms of 64 frequency bands and 100 frames per second (so it produces ~3 predictions per second, taking 2.84 seconds of context into account for a prediction, and its predictions still have 2 frequency bands that will be taken care of in the global pooling step). The receptive field of 284 frequency bands on a spectrogram of only 64 bands is achieved by excessive zero-padding (via padded convolutions).</p>\n<p>This model is much larger (~78 mio parameters) and slower to train, but achieves notably better results. This is in part since it is pretrained on AudioSet (training from scratch instead of using the pretrained weights reduced F1 score on the validation set from 0.688 to 0.676 in an early experiment).</p>\n<h2>Global pooling</h2>\n<p>Up to here, our model produces a time series of logits for each class. The final step is to pool these logits into a single prediction per class for the full recording, such that we can compute (and minimize) the classification error wrt. the given global labels for the recording.</p>\n<p>Reproducing the reasoning from my <a href=\"https://github.com/f0k/birdclef2018\" target=\"_blank\">BirdCLEF 2018 model</a>, there are two obvious ways and a third one in between:</p>\n<ul>\n<li><strong>mean pooling:</strong> This computes the average of the predictions over time, separately per class. The effect is that our prediction will be more confident for a bird that is detected many times during the recording compared to a bird that is detected once (even if very confidently).<br>\nThis is a bad idea: the label for a recording should reflect whether a bird is present, irrespective of how often it can be heard.<br>\nMean pooling will also distribute the gradient of the loss uniformly over all time points, training the network to predict each labeled species everywhere in the recording.</li>\n<li><strong>max pooling:</strong> This computes the maximum over time per class, so a single confident local detection will result in a confident global detection, matching the meaning of the global annotations.<br>\nHowever, max pooling will backpropagate the gradient of the loss only to the single time point where the prediction was most confident (separately for each species), pulling it up if the bird appears among the global labels, and pushing it down otherwise. This means a lot of computation for a very sparse update which ignores all the other vocalizations of the bird in the same recording.</li>\n<li><strong>log-mean-exp pooling:</strong> As a compromise, \\(\\frac{1}{a} \\log \\left( \\frac{1}{T} \\sum_{t=1}^{T} \\exp (a x_t) \\right) \\) allows to interpolate between taking the maximum (\\(a \\rightarrow \\infty\\)) and mean (\\(a \\rightarrow 0\\)). With \\(a=1\\), the output depends on the largest couple of values, which is also where the gradient is distributed to. This is what I used for most models. I also tried learning it, and learning a separate \\(a\\) per species (since some species might vocalize densely, warranting a small \\(a\\), and others sparsely, requiring a large \\(a\\)), but this only improved my validation scores without helping on the challenge data.</li>\n</ul>\n<p>Note that pooling operates in the logit domain here, that is, before applying the sigmoid that turns predictions into probabilities. This does not make a difference in max-pooling, but otherwise I empirically found that this improves results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121631,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "12/21/2020 19:25:16",
          "content": "<p>Thanks for the thorough writeup! </p>\n<p>I'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.</p>\n<p>(Also, the <a href=\"url\" target=\"_blank\">http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf</a> is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144820,
          "author_name": "janschl",
          "author_url": "",
          "post_date": "01/08/2021 17:32:37",
          "content": "<blockquote>\n  <p>I'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.</p>\n</blockquote>\n<p>I think I tried that some time and it produced worse results, but that was on a different dataset. Maybe I should try again, I'm curious now. I'll post here if I do.</p>\n<blockquote>\n  <p>(Also, the <a href=\"http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf\" target=\"_blank\">http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf</a> is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)</p>\n</blockquote>\n<p>Yes, I'm aware of it, but it's a good pointer! I only learned about it after BirdCLEF 2018. Their softmax-based pooling is highly related to log-mean-exp pooling, the latter was just around longer (Equation 6 in <a href=\"http://openaccess.thecvf.com/content_cvpr_2015/html/Pinheiro_From_Image-Level_to_2015_CVPR_paper.html\" target=\"_blank\">Pinheiro et al., 2015</a>).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121413,
      "author_name": "janschl",
      "author_url": "",
      "post_date": "12/21/2020 16:18:58",
      "content": "<h2>Data</h2>\n<p>We used both the official training data and <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">Rohan Rao's xeno-canto crawls</a>, totalling in 43499 recordings for the 200 species. We reserved about 10% for validation (splitting such that no recordist is part of both the training and validation set). Labels contain both the annotated foreground and background species for a recording, without distinguishing them (setting lower target values or lower weights for background species did not improve results).</p>\n<p>All recordings are predecoded to 22050 Hz 16-bit stereo or mono wave files. As part of the data loading pipeline, stereo files are <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L259\" target=\"_blank\">downmixed to mono</a> with a randomly uniform weight \\(p\\) for the left channel, and \\(1-p\\) for the right channel. This augmentation slightly improves results over downmixing them with fixed weights during decoding.</p>\n<p>To make the models work under low signal-to-noise ratios (i.e., the conditions found in the test files), we mix them with excerpts from the <a href=\"https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591\" target=\"_blank\">Chernobyl BiVA</a> and <a href=\"https://zenodo.org/record/1205569\" target=\"_blank\">BirdVox-full-night</a> datasets, which Paul suggested and hunted down. These datasets are precisely annotated with bird occurrences, so we can extract all parts void of birds. Mixing is <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L195\" target=\"_blank\">done on-the-fly</a> as part of the data loader. I started out carefully, but the best setting turned out to be mixing <em>every</em> training example with background noise, drawing a value \\(p \\in [0,1)\\) and scaling the noise with \\(p\\) and the bird recording excerpt with \\(1-p\\). For two of the models in the final ensemble, I also set \\(p=1\\) with 1% probability, setting the labels to all zero in this case.</p>\n<p>For model selection (explained further below), we also made use of the two official <code>example_test_audio</code> recordings, as well as the six 10-minute North American BirdCLEF 2020 validation set recordings <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911091\" target=\"_blank\">posted on the forum</a>. However, they only contain a fraction of the relevant species, contain species not part of the competition, and are not accurately labeled (missing several calls).</p>\n<h2>Training</h2>\n<p>Ideally, the model would be trained on complete recordings – this is the only way we can be sure all the targets are correct. If we pick a random excerpt, it is not guaranteed that all birds annotated to be present in the recording are also audible in the chosen excerpt. However, the longest recordings are 3 hours, which is impractical. As for BirdCLEF 2018, we train on randomly selected 30-second excerpts instead, hoping that most annotated birds will be audible at least once. Too short files are looped to make up 30 seconds (for BirdCLEF, we selected mini-batches among files of similar length to avoid unneeded looping, but PyTorch runs out of memory if the input tensors change shapes all the time). Validation uses the central 30 seconds of a recording.</p>\n<p>Training uses ADAM with mini-batches of 16 examples, an initial learning rate of 1e-3, and PyTorch's default settings for beta1, beta2 and epsilon. The validation loss is computed every 1000 update steps. If it does not improve over the current best value for 10 such evaluations in a row, the learning rate is reduced to a tenth, and training is continued. Training is stopped when the learning rate reaches 1e-6.</p>\n<p>For the PANN, I tried reducing the learning rate for the pretrained layers to 1% or 10% compared to the novel layers or to freeze the pretrained layers for some time, but it turned out that using the full learning rate for all the layers from the start works best.</p>\n<p>Training takes about 6h for a vanilla model, 7h for a residual network, and 8h for a PANN, on an RTX 2080 Ti.</p>\n<h2>Inference</h2>\n<p>The challenge required two types of inference:</p>\n<ol>\n<li>Predicting a set of species for every non-overlapping 5-second window of a 10-minute recording (for recording sites 1 and 2)</li>\n<li>Predicting a set of species for a full 10-minute recording (for recording site 3)</li>\n</ol>\n<p>Since the network is built to produce local predictions that are then pooled over time, we can simply pass a complete recording through the network up to the pooling operation, and then either</p>\n<ol>\n<li>Apply pooling with non-overlapping 5-second windows</li>\n<li>Apply pooling over the full recording</li>\n</ol>\n<p>The first case requires some trickery, since the network produces an uneven number of predictions within 5 seconds (5x70/9=38.889 for the vanilla model and ResNet, and 5x100/32=15.625 for the PANN). I solved it by having the models compute and report their receptive field (size, padding and stride) wrt. the audio samples, using this to <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L252-L258\" target=\"_blank\">compute a time point for each prediction</a>, and then <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L87\" target=\"_blank\">placing the pooling boundaries</a> appropriately. Note that this produces similar results as presegmenting the audio into 5-second windows and passing them through the network, but without artifacts introduced by zero-padding for the PANN, and with considerably less computation.</p>\n<p>Unfortunately, both types of inference do not match what the model was trained for: to produce good labels for 30-second audio excerpts. Since we use global log-mean-exp pooling, the output of the pooling stage is not completely independent of the input length; it is still somewhat resembling a global mean pooling:</p>\n<ol>\n<li>A short false positive that would stay under the threshold in a 30-second excerpt may score above the threshold for a 5-second window.</li>\n<li>If we feed 10-minute recordings, birds that are locally detected in some parts of the recording may still fall under the threshold after pooling.</li>\n</ol>\n<p>After 8 days of optimizing the models (4.5 days before the end of the challenge), I changed inference for the second case:</p>\n<p>&nbsp; &nbsp; 2. Apply pooling with 20-second windows overlapping by 50%, then take the maximum over the pooled windows for each species.</p>\n<p>Using 20 seconds instead of 30 seconds seemed to produce slightly better results on the validation data. However, this new scheme did not change my score on the public leaderboard – probably because it only affects the recordings from site 3, which may have a small influence on the total score (depending on how they are weighted). (On the private set, the score improved from 0.656 to 0.657, also unimportant.)</p>\n<p>After some more experimentation with models (improving the score to 0.587 by switching to PANNs), one day and 15 minutes before the deadline, I also changed inference for the first case (recording sites 1 and 2):</p>\n<ol>\n<li>First predict species globally for the 10-minute recording, using the new inference scheme developed for site 3. Then predict species locally by pooling in non-overlapping 5-second windows, but <strong>using a detection threshold of 0.3 instead of 0.5</strong>, and <strong>only keeping species that were predicted globally</strong> for that recording.</li>\n</ol>\n<p>The idea is that we should use longer windows to reliably detect birds, but once we are sure they are present in a recording, we can increase the classifier's sensitivity to find all occurrences.<br>\nThis gave a real boost from 0.587 to 0.600 on the public leaderboard. I thought about spending the second submission of that day on reducing the threshold to 0.2, but tried submitting an improved ensemble instead, which then finished uploading 30 seconds too late.</p>\n<p>Regarding computational efficiency, the final ensemble consisting of 3 PANNs took 4 minutes to run over the test set on Kaggle (but I do not know if and by how much the submission notebook evaluation is parallellized over subsets of the data).</p>\n<h2>Model selection</h2>\n<p>Most of our submissions used ensembles of 3 to 5 models selected by their F1 score on the validation set (10% of the training data from xeno-canto). The validation set may not be a good indicator for the challenge test set performance since a) it consists of focal recordings, not soundscapes, and b) it only comes with file-wise labels, not 5-second chunks or smaller. In addition, c) its distribution of species probably differs significantly from the test data.</p>\n<p>In addition to the validation set, we used the 8 annotated soundscape excerpts (2 from the challenge dataset, 6 from BirdCLEF 2020) to verify aspects such as the inference strategy. The 8 soundscape excerpts are also not a good indicator for the challenge performance since a) they only cover a small subset of the species and b) they are not thoroughly annotated.</p>\n<p>On the final day, I tried to improve model selection by setting up 6 different tasks:</p>\n<ol>\n<li>Prediction in 5-second windows on the 8 annotated soundscape excerpts</li>\n<li>Predicting file-level labels for the 8 annotated soundscape excerpts</li>\n<li>Predicting file-level labels on the validation set</li>\n<li>As 3., but recordings augmented with background noise from Chernobyl and BirdVox</li>\n<li>Predicting file-level labels only on those files from the validation set which have annotated background species (the other files may not be annotated completely)</li>\n<li>As 5., but recordings augmented with background noise from Chernobyl and BirdVox</li>\n</ol>\n<p>I evaluated all 16 previous submissions on these 6 tasks and looked for some correlation of precision, recall or F1 score with the submissions' public leaderboard performance:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F50bc581848f305ec9036bbca2313efed%2Fwriteup_correlation.png?generation=1608567691969180&amp;alt=media\"><br>\nThere was no clear indicator. With a combination of two scores that was at least almost monotonically increasing (when ordered by public leaderboard scores), I selected an ensemble that improved public leaderboard score from 0.600 to 0.601.</p>\n<p>Disappointed, for the last submission I chose another ensemble based on validation set performance alone. I hesitated whether to try reducing the prediction threshold from 0.3 to 0.2, but was afraid to ruin the submission. It scored 0.596 on the public leaderboard (it would have made the 4th place on the private set).</p>\n<h2>What did not work</h2>\n<ul>\n<li>Khaled ran several experiments using his <a href=\"https://arxiv.org/abs/1909.02859\" target=\"_blank\">receptive-field-regularized CNNs</a> which he successfully employed in past DCASE challenges, but unfortunately was not able to surpass the other models.</li>\n<li>Lukas contributed an <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L430\" target=\"_blank\">augmentation using colored noise</a>, but it turned out not to improve results over augmenting with the Chernobyl and BirdVox recordings. In addition, he worked on pretraining a model with self-supervised learning.</li>\n<li>Using shorter excerpts (and larger batches) made results worse, longer excerpts (and shorter batches) did not improve results.</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L405-L414\" target=\"_blank\">Reducing the frequency range</a> in order to have higher frequency resolution did not improve results.</li>\n<li>Augmentation via a cheap pitch shift implemented by <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L297-L303\" target=\"_blank\">randomly warping the mel filterbank</a> did not improve results.</li>\n<li>Training <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L272-L278\" target=\"_blank\">only on high-quality recordings</a>, or <a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L69-L74\" target=\"_blank\">weighting the loss by the quality rating</a> of the recording made results worse.</li>\n<li>Mixup made results worse (vanilla model, validation set results: 0.661 without mixup, 0.653 with alpha=0.1, 0.643 with alpha=0.3).</li>\n<li><a href=\"https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/metrics/__init__.py#L171-L174\" target=\"_blank\">Label smoothing</a> did not have a consistent effect (PANN, validation set results: 0.690 without, 0.682 with 1% smoothing, 0.687 with 5% smoothing).</li>\n<li><a href=\"http://proceedings.mlr.press/v80/furlanello18a.html\" target=\"_blank\">Born-Again Neural Networks</a> helped for BirdCLEF 2018, but not here, at least not for the PANN (validation set results: 0.690 originally, 0.685 as a born-again network).</li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>Apart from the challenge hosts, we are indebted to the following people:</p>\n<ul>\n<li>Hidehisa Arai for his <a href=\"https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline/#Prediction-loop\" target=\"_blank\">template on computing predictions</a> that falls back to a replacement test dataset at development time, when the actual test data is not available</li>\n<li>Alex Shonenkov for compiling this <a href=\"https://www.kaggle.com/shonenkov/birdcall-check\" target=\"_blank\">replacement test dataset</a></li>\n<li>Rohan Rao for crawling xeno-canto to compile his <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">additional training dataset</a></li>\n<li>All competitors who raised the bar on the leaderboard, and everyone discussing on the forum!</li>\n</ul>\n<h2>Afterthoughts</h2>\n<p>After the challenge finished, I was anxious to check if reducing the threshold from 0.3 to 0.2 would have made a difference. Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions 😃<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F70b5a32bbc3cd1b741c6696251c54c25%2Fmoneyknob.png?generation=1600330942624036&amp;alt=media\" alt=\"\"><br>\nA threshold of 0.2 would have ended on the third place, a threshold of 0.1 or less (even 0.02) would have made the second place.</p>\n<p><em>Were my submissions close to being in the money?</em><br>\nArguably yes, the models were fine, it only took changing a single hyperparameter in the inference pipeline that I was fully aware of. But as CPMP mentioned, that's probably true of other submissions as well!</p>\n<p><em>Was I close to submitting a winning solution?</em><br>\nThis thought has given me some unrest after the competition, but I think I wasn't.<br>\nAfter my first try of the new inference strategy, there were only three submissions remaining. I thought about trying a reduced threshold, but reducing the threshold in the old inference strategy had yielded worse results (despite an improved F-score on the 8 annotated soundscapes available for testing), so I was pessimistic about this. If I had tried anyway, I probably wouldn't have spent the last two submissions on reducing it far enough to be in the money with that ensemble.<br>\nOn the last submission, minutes before the deadline, I considered submitting the last new ensemble with threshold 0.2, but was afraid to change two things at the same time and ruin it. It would have yielded 0.599 on the public set. I would have thought I ruined it and not have selected it for the final standing.<br>\nIf I had started the whole competition earlier, I might have fallen for optimizing too much on the public set and missing the top ten. So overall I'm glad how it played out!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1014659,
      "author_name": "denisderonjic",
      "author_url": "",
      "post_date": "09/17/2020 16:19:19",
      "content": "<p>Congratulations!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1014127": "First of all, congratulations to all the participants, and thanks a lot to the organizers for this competition! It was great fun!\n\nIn short, our approach was based on [Jan's Sound Event Detection model from BirdCLEF 2018](https://github.com/f0k/birdclef2018), replacing the predictor with a pretrained CNN from [Qiuqiang Kong's PANN repository](https://github.com/qiuqiangkong/audioset_tagging_cnn). It was trained on 30-second snippets from the official training set extended with [Rohan Rao's xeno-canto crawls](https://www.kaggle.com/c/birdsong-recognition/discussion/159970), using binary cross-entropy against all labeled foreground and background species. Data was augmented by choosing random weights for downmixing the left and right channel to mono (for stereo files), and by mixing in bird-free background noise from the [Chernobyl BiVA](https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591) and [BirdVox-full-night](https://zenodo.org/record/1205569) datasets. Inference followed a two-stage procedure that first established a set of species for the recording, using 20-second windows and a threshold of 0.5, then looked for these species in 5-second windows with a threshold of 0.3. Code is [available on github](https://github.com/f0k/kagglebirds2020).\n\nThe following sections will explain things in more detail, spread over additional posts due to Kaggle's post size limit. Choose \"Sort by: Oldest\" to see them in their original order.\n\nIf you prefer a slide deck over text, feel free to click through [some slides explaining what we did](https://docs.google.com/presentation/d/10m0W13sJozYmfWPcPlaBg1j-MOW7p6w_-Tp3pG5KGE0/edit?usp=sharing).",
    "1014359": "> Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions\n\nUnfortunately this is true for many teams in every competition.  Looking for your writeup, and congrats on the result!",
    "1014384": "Congratulations 🎊 \n\nHope for next time.... Waiting for your writeup",
    "1014659": "Congratulations!!!",
    "1121403": "Team\n----\n\nWe are three PhD students (Khaled Koutini, Lukas Martak, Paul Primus) and a postdoc (Jan Schlüter) from the [Institute of Computational Perception](https://https://www.jku.at/en/institute-of-computational-perception/) at Johannes-Kepler-University Linz, Austria. Since 20 of our 23 submissions were done by me (Jan), most of this is written from my perspective, but each of us did important contributions I will highlight below.\n\n\nNetwork architecture\n--------------------\n\nThe general model architecture looks as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2Faf6861f1136591254f287a776e3682f1%2Fwriteup_network.png?generation=1608311897835230&alt=media)\n\n\nFrom an (arbitrarily long) monophonic raw audio recording, the part denoted as \"Frontend\" computes a spectrogram-like representation. In the next step, a Fully-Convolutional Network (FCN) processes this representation into a time series of logits for every class. When passed through a sigmoid, these would give us local predictions at every time step. Since we do not have local labels to train these, only file-wise labels, we apply a global pooling operation (over time) to obtain a single logit per class. Passed through a sigmoid, these serve as our file-wise predictions.\n\nThis layout matches what we used for [BirdCLEF 2018](https://github.com/f0k/birdclef2018), except that we are using a sigmoid instead of a softmax now (in BirdCLEF, evaluation was based on the ranking of species probabilities, for which a softmax helped). It also matches [what Hidehisa Arai explained](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection) early on in the competition.\n\nWe still have several options for the three components: The frontend, the local predictor, and the global pooling operation.\n\n### Frontend\n\n<details>\n<summary>PyTorch architecture listing:</summary>\n\n```\n(frontend): Sequential(\n  (filterbank): Sequential(\n    (stft): STFT(winsize=1024, hopsize=315, complex=False)\n    (melfilter): MelFilter(num_bands=80, min_freq=27.5, max_freq=10000.0)\n  )\n  (magscale): Log1p(trainable=True)\n  (denoise): SubtractMedian()\n  (norm): TemporalBatchNorm(\n    (bn): BatchNorm1d(80, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n  )\n)\n```\n</details>\n\nThe frontend takes in audio recordings at a sample rate of 22050 Hz. (While the test recordings were at 32000 Hz, not all training recordings go that high, so I resampled everything to 22050 Hz.)\n\nIt consists of:\n- [STFT](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L171):\n  window size of 1024 samples, hop size of 315 samples (resulting in 22050/315=70 frames per second), with Hann window, keeping only the magnitudes.\n- [mel filterbank](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L268):\n  80 triangular filters from 27.5 Hz to 10 kHz (not going up to 11025 Hz on purpose to leave some room for pitch shifting)\n- [nonlinear magnitude scaling](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L72):\n  Magnitudes are compressed by passing them through \\\\(y = \\log(1 + 10^a x)\\\\), where \\\\(a\\\\) is initialized to zero and learned by backpropagation.\n  I got similar, but maybe slightly worse results with \\\\(y = x^{\\sigma(a)}\\\\), where \\\\(a\\\\) is initialized to zero (resulting in \\\\(y = \\sqrt{x}\\\\)) and learned by backpropagation.\n- [denoising](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L163):\n  The different recordings have very different background noise floors, both due to the different environments and due to different recording equipment. Human listeners are quite good at adapting to this within a few seconds. To make it easier for the network, I unify the recordings somewhat by subtracting the median over time from each frequency band (separately for each recording or excerpt). I also tried [Per-Channel Energy Normalization](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L114), but subtracting the median worked better (and is faster). For PANNs this step is skipped, as it wouldn't match the input they are pretrained on.\n- [normalization](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L53):\n  To ensure inputs are in a reasonable range (and stay in a reasonable range when the magnitude scaling changes during training), each frequency band is normalized over time with Batch Normalization.\n\n### Local predictor\n\nThe purpose of the local predictor is to take the spectrogram produced by the frontend, and produce 200 time series of logits, one for each bird species.\nThe spectrogram can be regarded as an 80 pixel high one-channel image, and the output as a 1 pixel high 200-channel image. So what we need in between is a series of convolutions and pooling operations (i.e., a fully-convolutional network) that reduces the image height from 80 to 1, and produces 200 channels.\nFor illustrative purposes, assume we use a single convolutional layer for that: It would need 200 filters of height 80, but we are still free to choose the width (corresponding to the temporal context used for a single local prediction), and the horizontal stride (corresponding to the temporal density of predictions, i.e., we may find that we do not need predictions at the same rate as our 70 spectrogram frames per second).\nOf course, we went deeper than a single layer, but we can still compute the receptive field width (the temporal context used for a prediction) and the horizontal stride (the prediction rate) for a fully-convolutional network.\n\nI experimented with three different architectures: A \"vanilla\" ConvNet, a small residual network, and a pretrained PANN.\n\n### Vanilla\n\nThe most simple architecture (based on \"sparrow\" from a [2017 EUSIPCO paper of ours](http://ofai.at/~jan.schlueter/pubs/2017_eusipco.pdf)) was [specified in our code](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/defaults.vars#L21) as follows:\n```\nconv2d:64@3x3,bn2d,lrelu,\nconv2d:64@3x3,bn2d,\npool2d:max@3x3,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@3x3,bn2d,lrelu,\nconv2d:128@17x3,bn2d,\npool2d:max@5x3,lrelu,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\n```\nWhere\n* `conv2d:F@HxW` denotes an unpadded 2d convolution of `F` filters with height `H` (frequency bands) and width `W` (time frames),\n* `pool2d:max@HxW` denotes non-overlapping unpadded 2d max pooling of height `H` and width `W`,\n* `bn2d` denotes batch normalization,\n* `lrelu` denotes the leaky rectifier (LReLU) of leakiness 0.01,\n* and `dropout:0.5` denotes ordinary (non-spatial) dropout of 50%.\n\nThe last convolution is indicated with `C` channels, this will be replaced by 200, the number of classes. (The last two convolutions can be seen as two fully-connected layers of 1024 and 200 dimensions operating on each time step.)\n\nIn total, this has a receptive field of 79x103 (79 frequency bands, 103 spectrogram frames, about 1.5 seconds) and a temporal stride of 9 (~7.778 predictions per second). (Late insight: We will actually lose the highest frequency band due to the first pooling operation &ndash; we could have used a mel filterbank of 79 filters up to 9646.4 Hz instead for this model.) It clocks in at 3.6 mio parameters.\n\nAs our batch size was not that large (see the section on training), I later replaced batch normalization with group normalization, using 16 groups throughout the model. This improved performance a bit.\n\n### ResNet\n\nAs in [BirdCLEF 2018](https://github.com/f0k/birdclef2018), I extended the previous model into a simple residual network by replacing each of the first four convolutions with a residual block. Taken [from our config file](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L170):\n```\nadd[conv2d:64@3x3,bn2d,relu,\n    conv2d:64@3x3\n   |crop2d:2,conv2d:64@1x1],\nadd[bn2d,relu,conv2d:64@3x3,\n    bn2d,relu,conv2d:64@3x3\n   |crop2d:2],\npool2d:max@3x3,\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2,conv2d:128@1x1],\nadd[bn2d,relu,conv2d:128@3x3,\n    bn2d,relu,conv2d:128@3x3\n   |crop2d:2],\nbn2d,relu,\nconv2d:128@12x3,bn2d,\nlrelu,pool2d:max@5x3,\nconv2d:1024@1x9,bn2d,lrelu,\ndropout:0.5,conv2d:1024@1x1,bn2d,lrelu,\ndropout:0.5,conv2d:C@1x1\"\n```\nWhere:\n* `add[A|B]` adds the output of two branches, and\n* `crop2d:K` reduces the height and width by `K` pixels on each side, needed to match what is lost in the unpadded convolutions.\n\nThe receptive field is at 80x119, only minimally larger, and the temporal stride stays at 9 frames. With 3.73 mio parameters it is a little bit larger than the vanilla model, it trains longer, and performs a bit better.\n\nAgain, it later turned out this worked better with group normalization.\n\n### PANN\n\nFor this challenge, it was explicitly allowed to use models pretrained on external data, as long as they were disclosed [on the forum](https://www.kaggle.com/c/birdsong-recognition/discussion/158877). Among the models listed there is [Qiuqiang Kong's PANN repository](https://github.com/qiuqiangkong/audioset_tagging_cnn) of CNNs trained on AudioSet, which is also used in [Hidehi Saarai's notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection).\n\nI downloaded [the weights](https://zenodo.org/record/3987831) for his [Cnn14_16k](https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2542) model and [reproduced its central part](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L58) to stack it onto my frontend.\n\nThis required some adaptations to the frontend to produce spectrograms compatible with what the PANN expects. I could mostly adopt [the settings from the PANN repository](https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/542c8c8/pytorch/models.py#L2547-L2552), but as my inputs are sampled at 22050 Hz instead of 16000 Hz, I had to [scale up the FFT window and hop size accordingly](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L92-L94). In addition, I needed to [scale and shift](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/pann/__init__.py#L108) the result of my magnitude compression. I verified that remaining differences in the formulation of the mel filterbank did not affect predictions on an AudioSet example.\n\nOn top of Cnn14_16k`s 6 convolutional blocks (ending in 2048 feature maps; using only the first 4 or 5 blocks resulted in worse performance), I added two new 1x1 convolutions of 1024 and 200 channels each, with a leaky rectifier in between and 50% ordinary dropout applied to each.\n\nOverall, this model has a receptive field of 284x284 and a stride of 32x32, on spectrograms of 64 frequency bands and 100 frames per second (so it produces ~3 predictions per second, taking 2.84 seconds of context into account for a prediction, and its predictions still have 2 frequency bands that will be taken care of in the global pooling step). The receptive field of 284 frequency bands on a spectrogram of only 64 bands is achieved by excessive zero-padding (via padded convolutions).\n\nThis model is much larger (~78 mio parameters) and slower to train, but achieves notably better results. This is in part since it is pretrained on AudioSet (training from scratch instead of using the pretrained weights reduced F1 score on the validation set from 0.688 to 0.676 in an early experiment).\n\n## Global pooling\n\nUp to here, our model produces a time series of logits for each class. The final step is to pool these logits into a single prediction per class for the full recording, such that we can compute (and minimize) the classification error wrt. the given global labels for the recording.\n\nReproducing the reasoning from my [BirdCLEF 2018 model](https://github.com/f0k/birdclef2018), there are two obvious ways and a third one in between:\n* **mean pooling:** This computes the average of the predictions over time, separately per class. The effect is that our prediction will be more confident for a bird that is detected many times during the recording compared to a bird that is detected once (even if very confidently).\nThis is a bad idea: the label for a recording should reflect whether a bird is present, irrespective of how often it can be heard.\nMean pooling will also distribute the gradient of the loss uniformly over all time points, training the network to predict each labeled species everywhere in the recording.\n* **max pooling:** This computes the maximum over time per class, so a single confident local detection will result in a confident global detection, matching the meaning of the global annotations.\nHowever, max pooling will backpropagate the gradient of the loss only to the single time point where the prediction was most confident (separately for each species), pulling it up if the bird appears among the global labels, and pushing it down otherwise. This means a lot of computation for a very sparse update which ignores all the other vocalizations of the bird in the same recording.\n* **log-mean-exp pooling:** As a compromise, \\\\(\\frac{1}{a} \\log \\left( \\frac{1}{T} \\sum_{t=1}^{T} \\exp (a x_t) \\right) \\\\) allows to interpolate between taking the maximum (\\\\(a \\rightarrow \\infty\\\\)) and mean (\\\\(a \\rightarrow 0\\\\)). With \\\\(a=1\\\\), the output depends on the largest couple of values, which is also where the gradient is distributed to. This is what I used for most models. I also tried learning it, and learning a separate \\\\(a\\\\) per species (since some species might vocalize densely, warranting a small \\\\(a\\\\), and others sparsely, requiring a large \\\\(a\\\\)), but this only improved my validation scores without helping on the challenge data.\n\nNote that pooling operates in the logit domain here, that is, before applying the sigmoid that turns predictions into probabilities. This does not make a difference in max-pooling, but otherwise I empirically found that this improves results.",
    "1121413": "Data\n----\n\nWe used both the official training data and [Rohan Rao's xeno-canto crawls](https://www.kaggle.com/c/birdsong-recognition/discussion/159970), totalling in 43499 recordings for the 200 species. We reserved about 10% for validation (splitting such that no recordist is part of both the training and validation set). Labels contain both the annotated foreground and background species for a recording, without distinguishing them (setting lower target values or lower weights for background species did not improve results).\n\nAll recordings are predecoded to 22050 Hz 16-bit stereo or mono wave files. As part of the data loading pipeline, stereo files are [downmixed to mono](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L259) with a randomly uniform weight \\\\(p\\\\) for the left channel, and \\\\(1-p\\\\) for the right channel. This augmentation slightly improves results over downmixing them with fixed weights during decoding.\n\nTo make the models work under low signal-to-noise ratios (i.e., the conditions found in the test files), we mix them with excerpts from the [Chernobyl BiVA](https://doi.org/10.5285/be5639e9-75e9-4aa3-afdd-65ba80352591) and [BirdVox-full-night](https://zenodo.org/record/1205569) datasets, which Paul suggested and hunted down. These datasets are precisely annotated with bird occurrences, so we can extract all parts void of birds. Mixing is [done on-the-fly](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L195) as part of the data loader. I started out carefully, but the best setting turned out to be mixing *every* training example with background noise, drawing a value \\\\(p \\in [0,1)\\\\) and scaling the noise with \\\\(p\\\\) and the bird recording excerpt with \\\\(1-p\\\\). For two of the models in the final ensemble, I also set \\\\(p=1\\\\) with 1% probability, setting the labels to all zero in this case.\n\nFor model selection (explained further below), we also made use of the two official `example_test_audio` recordings, as well as the six 10-minute North American BirdCLEF 2020 validation set recordings [posted on the forum](https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911091). However, they only contain a fraction of the relevant species, contain species not part of the competition, and are not accurately labeled (missing several calls).\n\n\nTraining\n--------\n\nIdeally, the model would be trained on complete recordings &ndash; this is the only way we can be sure all the targets are correct. If we pick a random excerpt, it is not guaranteed that all birds annotated to be present in the recording are also audible in the chosen excerpt. However, the longest recordings are 3 hours, which is impractical. As for BirdCLEF 2018, we train on randomly selected 30-second excerpts instead, hoping that most annotated birds will be audible at least once. Too short files are looped to make up 30 seconds (for BirdCLEF, we selected mini-batches among files of similar length to avoid unneeded looping, but PyTorch runs out of memory if the input tensors change shapes all the time). Validation uses the central 30 seconds of a recording.\n\nTraining uses ADAM with mini-batches of 16 examples, an initial learning rate of 1e-3, and PyTorch's default settings for beta1, beta2 and epsilon. The validation loss is computed every 1000 update steps. If it does not improve over the current best value for 10 such evaluations in a row, the learning rate is reduced to a tenth, and training is continued. Training is stopped when the learning rate reaches 1e-6.\n\nFor the PANN, I tried reducing the learning rate for the pretrained layers to 1% or 10% compared to the novel layers or to freeze the pretrained layers for some time, but it turned out that using the full learning rate for all the layers from the start works best.\n\nTraining takes about 6h for a vanilla model, 7h for a residual network, and 8h for a PANN, on an RTX 2080 Ti.\n\n\nInference\n---------\n\nThe challenge required two types of inference:\n1. Predicting a set of species for every non-overlapping 5-second window of a 10-minute recording (for recording sites 1 and 2)\n2. Predicting a set of species for a full 10-minute recording (for recording site 3)\n\nSince the network is built to produce local predictions that are then pooled over time, we can simply pass a complete recording through the network up to the pooling operation, and then either\n1. Apply pooling with non-overlapping 5-second windows\n2. Apply pooling over the full recording\n\nThe first case requires some trickery, since the network produces an uneven number of predictions within 5 seconds (5x70/9=38.889 for the vanilla model and ResNet, and 5x100/32=15.625 for the PANN). I solved it by having the models compute and report their receptive field (size, padding and stride) wrt. the audio samples, using this to [compute a time point for each prediction](https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L252-L258), and then [placing the pooling boundaries](https://github.com/CPJKU/kagglebirds2020/blob/master/predict_kagglebirds2020.py#L87) appropriately. Note that this produces similar results as presegmenting the audio into 5-second windows and passing them through the network, but without artifacts introduced by zero-padding for the PANN, and with considerably less computation.\n\nUnfortunately, both types of inference do not match what the model was trained for: to produce good labels for 30-second audio excerpts. Since we use global log-mean-exp pooling, the output of the pooling stage is not completely independent of the input length; it is still somewhat resembling a global mean pooling:\n1. A short false positive that would stay under the threshold in a 30-second excerpt may score above the threshold for a 5-second window.\n2. If we feed 10-minute recordings, birds that are locally detected in some parts of the recording may still fall under the threshold after pooling.\n\nAfter 8 days of optimizing the models (4.5 days before the end of the challenge), I changed inference for the second case:\n\n&nbsp; &nbsp; 2. Apply pooling with 20-second windows overlapping by 50%, then take the maximum over the pooled windows for each species.\n\nUsing 20 seconds instead of 30 seconds seemed to produce slightly better results on the validation data. However, this new scheme did not change my score on the public leaderboard &ndash; probably because it only affects the recordings from site 3, which may have a small influence on the total score (depending on how they are weighted). (On the private set, the score improved from 0.656 to 0.657, also unimportant.)\n\nAfter some more experimentation with models (improving the score to 0.587 by switching to PANNs), one day and 15 minutes before the deadline, I also changed inference for the first case (recording sites 1 and 2):\n1. First predict species globally for the 10-minute recording, using the new inference scheme developed for site 3. Then predict species locally by pooling in non-overlapping 5-second windows, but **using a detection threshold of 0.3 instead of 0.5**, and **only keeping species that were predicted globally** for that recording.\n\nThe idea is that we should use longer windows to reliably detect birds, but once we are sure they are present in a recording, we can increase the classifier's sensitivity to find all occurrences.\nThis gave a real boost from 0.587 to 0.600 on the public leaderboard. I thought about spending the second submission of that day on reducing the threshold to 0.2, but tried submitting an improved ensemble instead, which then finished uploading 30 seconds too late.\n\nRegarding computational efficiency, the final ensemble consisting of 3 PANNs took 4 minutes to run over the test set on Kaggle (but I do not know if and by how much the submission notebook evaluation is parallellized over subsets of the data).\n\n\nModel selection\n---------------\n\nMost of our submissions used ensembles of 3 to 5 models selected by their F1 score on the validation set (10% of the training data from xeno-canto). The validation set may not be a good indicator for the challenge test set performance since a) it consists of focal recordings, not soundscapes, and b) it only comes with file-wise labels, not 5-second chunks or smaller. In addition, c) its distribution of species probably differs significantly from the test data.\n\nIn addition to the validation set, we used the 8 annotated soundscape excerpts (2 from the challenge dataset, 6 from BirdCLEF 2020) to verify aspects such as the inference strategy. The 8 soundscape excerpts are also not a good indicator for the challenge performance since a) they only cover a small subset of the species and b) they are not thoroughly annotated.\n\nOn the final day, I tried to improve model selection by setting up 6 different tasks:\n1. Prediction in 5-second windows on the 8 annotated soundscape excerpts\n2. Predicting file-level labels for the 8 annotated soundscape excerpts\n3. Predicting file-level labels on the validation set\n4. As 3., but recordings augmented with background noise from Chernobyl and BirdVox\n5. Predicting file-level labels only on those files from the validation set which have annotated background species (the other files may not be annotated completely)\n6. As 5., but recordings augmented with background noise from Chernobyl and BirdVox\n\nI evaluated all 16 previous submissions on these 6 tasks and looked for some correlation of precision, recall or F1 score with the submissions' public leaderboard performance:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F50bc581848f305ec9036bbca2313efed%2Fwriteup_correlation.png?generation=1608567691969180&alt=media\" width=\"300\">\nThere was no clear indicator. With a combination of two scores that was at least almost monotonically increasing (when ordered by public leaderboard scores), I selected an ensemble that improved public leaderboard score from 0.600 to 0.601.\n\nDisappointed, for the last submission I chose another ensemble based on validation set performance alone. I hesitated whether to try reducing the prediction threshold from 0.3 to 0.2, but was afraid to ruin the submission. It scored 0.596 on the public leaderboard (it would have made the 4th place on the private set).\n\n\nWhat did not work\n-----------------\n\n* Khaled ran several experiments using his [receptive-field-regularized CNNs](https://arxiv.org/abs/1909.02859) which he successfully employed in past DCASE challenges, but unfortunately was not able to surpass the other models.\n* Lukas contributed an [augmentation using colored noise](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/datasets/kagglebirds2020/__init__.py#L430), but it turned out not to improve results over augmenting with the Chernobyl and BirdVox recordings. In addition, he worked on pretraining a model with self-supervised learning.\n* Using shorter excerpts (and larger batches) made results worse, longer excerpts (and shorter batches) did not improve results.\n* [Reducing the frequency range](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L405-L414) in order to have higher frequency resolution did not improve results.\n* Augmentation via a cheap pitch shift implemented by [randomly warping the mel filterbank](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/models/audioclass/frontend.py#L297-L303) did not improve results.\n* Training [only on high-quality recordings](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L272-L278), or [weighting the loss by the quality rating](https://github.com/CPJKU/kagglebirds2020/blob/master/experiments/train_kagglebirds2020.sh#L69-L74) of the recording made results worse.\n* Mixup made results worse (vanilla model, validation set results: 0.661 without mixup, 0.653 with alpha=0.1, 0.643 with alpha=0.3).\n* [Label smoothing](https://github.com/CPJKU/kagglebirds2020/blob/master/definitions/metrics/__init__.py#L171-L174) did not have a consistent effect (PANN, validation set results: 0.690 without, 0.682 with 1% smoothing, 0.687 with 5% smoothing).\n* [Born-Again Neural Networks](http://proceedings.mlr.press/v80/furlanello18a.html) helped for BirdCLEF 2018, but not here, at least not for the PANN (validation set results: 0.690 originally, 0.685 as a born-again network).\n\n\nAcknowledgements\n----------------\n\nApart from the challenge hosts, we are indebted to the following people:\n* Hidehisa Arai for his [template on computing predictions](https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline/#Prediction-loop) that falls back to a replacement test dataset at development time, when the actual test data is not available\n* Alex Shonenkov for compiling this [replacement test dataset](https://www.kaggle.com/shonenkov/birdcall-check)\n* Rohan Rao for crawling xeno-canto to compile his [additional training dataset](https://www.kaggle.com/c/birdsong-recognition/discussion/159970)\n* All competitors who raised the bar on the leaderboard, and everyone discussing on the forum!\n\n\nAfterthoughts\n-------------\n\nAfter the challenge finished, I was anxious to check if reducing the threshold from 0.3 to 0.2 would have made a difference. Turns out I had a knob to make money that I didn't try due to a lack of remaining submissions 😃\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2546802%2F70b5a32bbc3cd1b741c6696251c54c25%2Fmoneyknob.png?generation=1600330942624036&alt=media)\nA threshold of 0.2 would have ended on the third place, a threshold of 0.1 or less (even 0.02) would have made the second place.\n\n*Were my submissions close to being in the money?*\nArguably yes, the models were fine, it only took changing a single hyperparameter in the inference pipeline that I was fully aware of. But as CPMP mentioned, that's probably true of other submissions as well!\n\n*Was I close to submitting a winning solution?*\nThis thought has given me some unrest after the competition, but I think I wasn't.\nAfter my first try of the new inference strategy, there were only three submissions remaining. I thought about trying a reduced threshold, but reducing the threshold in the old inference strategy had yielded worse results (despite an improved F-score on the 8 annotated soundscapes available for testing), so I was pessimistic about this. If I had tried anyway, I probably wouldn't have spent the last two submissions on reducing it far enough to be in the money with that ensemble.\nOn the last submission, minutes before the deadline, I considered submitting the last new ensemble with threshold 0.2, but was afraid to change two things at the same time and ruin it. It would have yielded 0.599 on the public set. I would have thought I ruined it and not have selected it for the final standing.\nIf I had started the whole competition earlier, I might have fallen for optimizing too much on the public set and missing the top ten. So overall I'm glad how it played out!",
    "1121631": "Thanks for the thorough writeup! \n\nI'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.\n\n(Also, the [http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf](url) is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)",
    "1144820": "> I'm curious if you tried pooling on the penultimate representations, rather than the logits - ie, pool, then apply a final dense layer.\n\nI think I tried that some time and it produced worse results, but that was on a different dataset. Maybe I should try again, I'm curious now. I'll post here if I do.\n\n> (Also, the http://www.justinsalamon.com/uploads/4/3/9/4/4394963/mcfee_autopool_taslp_2018.pdf is a nice point of reference for the 'inbetween' pooling, if you haven't seen it.)\n\nYes, I'm aware of it, but it's a good pointer! I only learned about it after BirdCLEF 2018. Their softmax-based pooling is highly related to log-mean-exp pooling, the latter was just around longer (Equation 6 in [Pinheiro et al., 2015](http://openaccess.thecvf.com/content_cvpr_2015/html/Pinheiro_From_Image-Level_to_2015_CVPR_paper.html))."
  },
  "source": "meta"
}