{
  "id": 220522,
  "title": "3rd Place Solution",
  "url": "/competitions/rfcx-species-audio-detection/writeups/sidchin-3rd-place-solution",
  "author_name": "",
  "post_date": "2021-02-23T22:05:24.900Z",
  "votes": 61,
  "comment_count": 22,
  "views": 0,
  "content": "<h1>3rd Place Solution</h1>\n<p><strong>TLDR</strong></p>\n<p>Our solution is a mean blend of 8 models trained on True positive labels of all recording ids in  train_tp.csv (given + hand-labeled labels) also from some recording ids in train_fp.csv (hand-labeled labels) with heavy augmentations. Additionally, some models are also trained on pseudo labels and a hand-labeled external dataset. We also post-processed the blended results by thresholding species 3.</p>\n<p>From 308 submissions it is obvious that we have tested a lot of techniques and I will share more detailed information by category below.</p>\n<h2>Data Preparation</h2>\n<p>I couldn't get a proper validation framework setup after trying out many techniques and decided at one point to start digging into the data and figured out that there are many unlabeled samples both in and out of the range of t_min and t_max labels given in train_tp.csv. In <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735</a> it was mentioned that using hand-labeled species was allowed, so I started labeling the data manually and after labeling 100 recording ids I could already get a &gt; 0.9 public lb score and pretty consistent local CV scores that somewhat correlates with the public lb. Naturally, I continued to label the entire train_tp.csv seeing that it has only around 1.3k recording ids.  Further labeling of train_fp.csv helped the score but only minimally so I stopped at one point. As I grew more familiar with the data I could label 300 recording ids in a day :), referring to pseudo labels helped a lot too. I also went through the train_tp.csv a few more rounds to make sure I have quality data. I used both spectrograms and listening strategy to analyze and label the data, some species are easy to spot with spectrograms and some are easier to spot by listening, and in some cases, both listening and visual inspection of the spectrograms can act as a multi verification technique to get more quality labels, especially when birds/frogs are very distant away from the recorder or there are strong noises like waterfall sounds. By labeling and analyzing the data I also figured out the kinds of sounds/noises that would appear and inspired me to try out a few augmentation methods which I will share below. Along with true positive labels, I also added noisy and non-noisy labels based on my confidence in the completeness of labels in a specific recording id. I am not a perfect labeler so I wanted to handle complete and non-complete labeled recording ids differently, which I will share below too.</p>\n<p>I also removed some labels from train_tp.csv as I found some true positives suspicious, I didn't test not removing the labels before so not sure how much this helped.</p>\n<p>Additionally, after finding out the paper from the organizers I searched for suitably licensed datasets with those species and found one dataset with species in this competition with a proper license. But there weren't any labels so I labeled it manually too with the same format as train_tp.csv. <a href=\"https://datadryad.org/stash/dataset/doi:10.5061/dryad.c0g2t\" target=\"_blank\">https://datadryad.org/stash/dataset/doi:10.5061/dryad.c0g2t</a> . I reuploaded the dataset <a href=\"https://www.kaggle.com/dicksonchin93/eleutherodactylus-frogs\" target=\"_blank\">HERE</a> with my manual labels.</p>\n<p>I uploaded the extra labels as a dataset <a href=\"https://www.kaggle.com/dicksonchin93/extra-labels-for-rcfx-competition-data\" target=\"_blank\">https://www.kaggle.com/dicksonchin93/extra-labels-for-rcfx-competition-data</a>, feel free to use it and see if you can get a better single model score! mine was 0.970 on public lb</p>\n<h2>Modeling / Data Pre-processing</h2>\n<p>I used Mel Spectrograms with the following parameters: 32kHz sampling rate, a hop size of 716, a window size of 1366, and 224 or 128 Number of Mels. Tried a bunch of methods but plainly using 3 layers of standardized Mel Spectrograms works the best. The image dimensions were (num_mel_bins, 750). </p>\n<p>Using train_tp.csv to create folds will potentially leak some training data into your validation data so I treated the problem as a multilabel target and used <a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a> to stratify the data into 5 partitions using unique recording ids and its multilabel targets. I had two different 5 fold partitions using different versions of the multi labels and used a mix of both in the final submission. </p>\n<p>I used multiple different audio duration during the competition and at different stages of the competition, the best duration varied in my implementation but in the end, I used 5 seconds of audio for training and prediction as the LWLWRAP score was better on both public lb and local validation. </p>\n<p>The 5-second audio was randomly sampled during training and in prediction time a 5-second sliding window was used with overlap and the max of predictions was used. How the 5-second audio is randomly sampled is considered to be an augmentation method in my opinion and so I will explain it in the heavy augmentations category below</p>\n<h6>Augmentations</h6>\n<ul>\n<li>Random 5-second audio samples: <br>\na starting point was chosen randomly on values between reference t_mins and t_maxes obtained from </li>\n</ul>\n<pre><code>def get_ref_tmin_tmax_and_species_ids(\n    self, all_tp_events, label_column_key=\"species_id\"\n):\n        all_tp_events[\"t_min_ref\"] = all_tp_events[\"t_min\"].apply(\n           lambda x: max(x - (self.period / 2.0), 0)\n        )\n        def get_tmax_ref(row, period=self.period):\n            tmin_x = row[\"t_min\"]\n            tmax_x = row[\"t_max\"]\n            tmax_ref = tmax_x - (period / 4.0)\n            if tmax_ref &lt; tmin_x:\n                tmax_ref = (tmax_x - tmin_x) / 2.0 + tmin_x\n            return tmax_ref\n        all_tp_events[\"t_max_ref\"] = all_tp_events[\n            [\"t_max\", \"t_min\"]\n        ].apply(get_tmax_ref, axis=1)\n        t_min_maxes = all_tp_events[\n            [\"t_min_ref\", \"t_max_ref\"]\n        ].values.tolist()\n        species_ids = all_tp_events[label_column_key].values.tolist()\n        return t_min_maxes, species_ids\n</code></pre>\n<p>Labels were also assigned based on the chosen starting time and ending time with t_min and t_max labels.</p>\n<ul>\n<li>audio based pink noise</li>\n<li>audio based white noise</li>\n<li>reverberation</li>\n<li>time stretch</li>\n<li>use one of 16kHz or 48kHz sample rate data and resample it to 32kHz sample rate using randomly chosen resampling methods <code>['kaiser_best', 'kaiser_fast', 'fft', 'polyphase']</code></li>\n<li>use different window types to compute spectrograms at train time <code>['flattop', 'hamming', ('kaiser', 4.0), 'blackman', 'hann']</code>  , hann window is used at test and validation time</li>\n<li>masking out non labeled chunks of the audio with a 10% chance</li>\n<li>one of spectrogram <a href=\"https://arxiv.org/abs/2002.12047\" target=\"_blank\">FMix</a> and audio based mixup with the max of labels instead of using the blend from the beta parameter</li>\n<li><a href=\"http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf\" target=\"_blank\">spec mix </a>: <br>\nonly one strip was used for each axis, for the horizontal axis when the chosen frequency range to mask out completely covers a specific species minimum f_min and maximum f_max , that species label will be dropped. Specmix is also using the max of labels instead of using the blend from the beta parameter. The code below shows how I obtain the function that can output frequency axis spectrogram positions from frequency</li>\n</ul>\n<pre><code>def get_mel_scaled_hz_to_y_axis_func(fmin=0, fmax=16000, n_mels=128):\n    hz_points = librosa.core.mel_frequencies(n_mels=n_mels, fmin=fmin, fmax=fmax)\n    hz_to_y_axis = interp1d(hz_points, np.arange(n_mels)[::-1])  \n    # reversed because first index is at the top left in an image array\n    return hz_to_y_axis\n</code></pre>\n<ul>\n<li>bandpass noise </li>\n<li>Water from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Engine and Motor Sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Honk, Traffic and Horn sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Speech sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Bark sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n</ul>\n<p>checkout recording_id b8d1e4865 to find dogs barking and some human speech :D</p>\n<h6>Architectures used</h6>\n<p>No SED just plain classifier models with GEM pooling for CNN based models</p>\n<ul>\n<li>Efficientnet-b7</li>\n<li>Efficientnet-b8</li>\n<li>HRNet w64</li>\n<li>deitbase224</li>\n<li>vit_large_patch16_224</li>\n<li>ecaresnet50</li>\n<li>2x resnest50 from <a href=\"https://www.kaggle.com/meaninglesslives\" target=\"_blank\">https://www.kaggle.com/meaninglesslives</a>, checkout his writeup in a minimal notebook <a href=\"https://www.kaggle.com/meaninglesslives/rfcx-minimal\" target=\"_blank\">HERE</a>!</li>\n</ul>\n<h6>Loss</h6>\n<p>The main loss strategy used for the final submission was using different loss function for samples which I am confident is complete in labels and samples which I am not confident is complete in labels. BCE was used for non-noisy/confident samples and a modified Lsoft loss was used on the noisy/non-confident. Lsoft loss was modified to be applied only to nonpositive samples, as I was confident in my manual labels. It looks like this </p>\n<pre><code>def l_soft_on_negative_samples(y_pred, y_true, beta, eps = 1e-7):\n    y_pred = torch.clamp(y_pred, eps, 1.0)\n\n    # (1) dynamically update the targets based on the current state of the model:\n    # bootstrapped target tensor\n    # use predicted class proba directly to generate regression targets\n    with torch.no_grad():\n        negative_indexes = (y_true == 0).nonzero().squeeze(1)\n        y_true_update = y_true\n        y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] = (\n            y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] * beta + \n            (1 - beta) * y_pred[negative_indexes[:, 0], negative_indexes[:, 1]]\n        )\n\n    # (2) compute loss as always\n    loss = F.binary_cross_entropy(y_pred, y_true_update)\n    return loss\n</code></pre>\n<p>This was inspired by the first placed winner in the Freesound competition <a href=\"https://github.com/lRomul/argus-freesound\" target=\"_blank\">https://github.com/lRomul/argus-freesound</a> but I noticed that it doesn't make sense if it is used with mixup since audio will be mixed up anyways. So I also obtain the max of noisy binary labels so that noisy labels mixed with clean labels are considered to be noisy labels.</p>\n<h6>Pseudo Labels</h6>\n<p>I didn't get much boost from pseudo labels, maybe I did something wrong but nonetheless, it was used in some models. I used a 0.8 threshold for labels generated with 5-second windows and utilized the same window positions during training. Using raw predictions didn't help the model at all on lb.</p>\n<h6>Post processing</h6>\n<p>We set the species 3 labels to be 1 with a 0.95 threshold and it boosted the score slightly</p>\n<h6>Other stuff</h6>\n<ul>\n<li>Early stopping of 20 epochs with a minimum learning rate of 9e-6 to start counting these 20 epochs</li>\n<li>Reduce learning rate on Plateau with a factor of 0.6 and start with a few warmup epochs, when LR is reduced the best model weights was loaded back again</li>\n<li>Catalyst was used</li>\n</ul>\n<h2>Things that failed</h2>\n<ul>\n<li>using models without pre-trained weights</li>\n<li>timeshift</li>\n<li>using species from the Cornell Competition that are confused with species in this competition as a distractor noise, for example, moudov is similar to species 15, reevir is similar to species 11, rebwoo is similar to species 6, bkbwar is similar to species 7, cacwre is similar to species 19, every is similar to species 17 and nrwswa is similar to species 20</li>\n<li>using plane sounds from Freesound50k data</li>\n<li>using PCEN, deltas or CQT</li>\n<li><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">Random Power</a></li>\n<li>TTA with different window types</li>\n<li>Manifold mixup with resnest50</li>\n<li>using trainable Switchnorm as an initial layer replacing normal standardization</li>\n<li>using trainable Examplar norm as an initial layer replacing normal standardization</li>\n<li>Context Gating</li>\n<li>split audio intro three equal-length chunks and concat as 3 layer image</li>\n<li>lsep and <a href=\"https://arxiv.org/abs/2009.14119\" target=\"_blank\">Assymetric loss </a></li>\n<li>using rain sounds from Freesound50k data</li>\n<li>using a fixed validation mask similar to how I used random training mask</li>\n<li>use <a href=\"https://vsitzmann.github.io/siren/\" target=\"_blank\">SIREN</a> layer</li>\n<li>Tried to separate some confusing patterns as separate manual labels but didn't get the chance to test them</li>\n</ul>\n<p>Hopefully, I didn't miss anything. Oh, we were holding off submitting a good model until <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> came along :)</p>",
  "messages": [
    {
      "id": "1209066",
      "postDate": "02/18/2021 16:59:07",
      "content": "<h1>3rd Place Solution</h1>\n<p><strong>TLDR</strong></p>\n<p>Our solution is a mean blend of 8 models trained on True positive labels of all recording ids in  train_tp.csv (given + hand-labeled labels) also from some recording ids in train_fp.csv (hand-labeled labels) with heavy augmentations. Additionally, some models are also trained on pseudo labels and a hand-labeled external dataset. We also post-processed the blended results by thresholding species 3.</p>\n<p>From 308 submissions it is obvious that we have tested a lot of techniques and I will share more detailed information by category below.</p>\n<h2>Data Preparation</h2>\n<p>I couldn't get a proper validation framework setup after trying out many techniques and decided at one point to start digging into the data and figured out that there are many unlabeled samples both in and out of the range of t_min and t_max labels given in train_tp.csv. In <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735</a> it was mentioned that using hand-labeled species was allowed, so I started labeling the data manually and after labeling 100 recording ids I could already get a &gt; 0.9 public lb score and pretty consistent local CV scores that somewhat correlates with the public lb. Naturally, I continued to label the entire train_tp.csv seeing that it has only around 1.3k recording ids.  Further labeling of train_fp.csv helped the score but only minimally so I stopped at one point. As I grew more familiar with the data I could label 300 recording ids in a day :), referring to pseudo labels helped a lot too. I also went through the train_tp.csv a few more rounds to make sure I have quality data. I used both spectrograms and listening strategy to analyze and label the data, some species are easy to spot with spectrograms and some are easier to spot by listening, and in some cases, both listening and visual inspection of the spectrograms can act as a multi verification technique to get more quality labels, especially when birds/frogs are very distant away from the recorder or there are strong noises like waterfall sounds. By labeling and analyzing the data I also figured out the kinds of sounds/noises that would appear and inspired me to try out a few augmentation methods which I will share below. Along with true positive labels, I also added noisy and non-noisy labels based on my confidence in the completeness of labels in a specific recording id. I am not a perfect labeler so I wanted to handle complete and non-complete labeled recording ids differently, which I will share below too.</p>\n<p>I also removed some labels from train_tp.csv as I found some true positives suspicious, I didn't test not removing the labels before so not sure how much this helped.</p>\n<p>Additionally, after finding out the paper from the organizers I searched for suitably licensed datasets with those species and found one dataset with species in this competition with a proper license. But there weren't any labels so I labeled it manually too with the same format as train_tp.csv. <a href=\"https://datadryad.org/stash/dataset/doi:10.5061/dryad.c0g2t\" target=\"_blank\">https://datadryad.org/stash/dataset/doi:10.5061/dryad.c0g2t</a> . I reuploaded the dataset <a href=\"https://www.kaggle.com/dicksonchin93/eleutherodactylus-frogs\" target=\"_blank\">HERE</a> with my manual labels.</p>\n<p>I uploaded the extra labels as a dataset <a href=\"https://www.kaggle.com/dicksonchin93/extra-labels-for-rcfx-competition-data\" target=\"_blank\">https://www.kaggle.com/dicksonchin93/extra-labels-for-rcfx-competition-data</a>, feel free to use it and see if you can get a better single model score! mine was 0.970 on public lb</p>\n<h2>Modeling / Data Pre-processing</h2>\n<p>I used Mel Spectrograms with the following parameters: 32kHz sampling rate, a hop size of 716, a window size of 1366, and 224 or 128 Number of Mels. Tried a bunch of methods but plainly using 3 layers of standardized Mel Spectrograms works the best. The image dimensions were (num_mel_bins, 750). </p>\n<p>Using train_tp.csv to create folds will potentially leak some training data into your validation data so I treated the problem as a multilabel target and used <a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a> to stratify the data into 5 partitions using unique recording ids and its multilabel targets. I had two different 5 fold partitions using different versions of the multi labels and used a mix of both in the final submission. </p>\n<p>I used multiple different audio duration during the competition and at different stages of the competition, the best duration varied in my implementation but in the end, I used 5 seconds of audio for training and prediction as the LWLWRAP score was better on both public lb and local validation. </p>\n<p>The 5-second audio was randomly sampled during training and in prediction time a 5-second sliding window was used with overlap and the max of predictions was used. How the 5-second audio is randomly sampled is considered to be an augmentation method in my opinion and so I will explain it in the heavy augmentations category below</p>\n<h6>Augmentations</h6>\n<ul>\n<li>Random 5-second audio samples: <br>\na starting point was chosen randomly on values between reference t_mins and t_maxes obtained from </li>\n</ul>\n<pre><code>def get_ref_tmin_tmax_and_species_ids(\n    self, all_tp_events, label_column_key=\"species_id\"\n):\n        all_tp_events[\"t_min_ref\"] = all_tp_events[\"t_min\"].apply(\n           lambda x: max(x - (self.period / 2.0), 0)\n        )\n        def get_tmax_ref(row, period=self.period):\n            tmin_x = row[\"t_min\"]\n            tmax_x = row[\"t_max\"]\n            tmax_ref = tmax_x - (period / 4.0)\n            if tmax_ref &lt; tmin_x:\n                tmax_ref = (tmax_x - tmin_x) / 2.0 + tmin_x\n            return tmax_ref\n        all_tp_events[\"t_max_ref\"] = all_tp_events[\n            [\"t_max\", \"t_min\"]\n        ].apply(get_tmax_ref, axis=1)\n        t_min_maxes = all_tp_events[\n            [\"t_min_ref\", \"t_max_ref\"]\n        ].values.tolist()\n        species_ids = all_tp_events[label_column_key].values.tolist()\n        return t_min_maxes, species_ids\n</code></pre>\n<p>Labels were also assigned based on the chosen starting time and ending time with t_min and t_max labels.</p>\n<ul>\n<li>audio based pink noise</li>\n<li>audio based white noise</li>\n<li>reverberation</li>\n<li>time stretch</li>\n<li>use one of 16kHz or 48kHz sample rate data and resample it to 32kHz sample rate using randomly chosen resampling methods <code>['kaiser_best', 'kaiser_fast', 'fft', 'polyphase']</code></li>\n<li>use different window types to compute spectrograms at train time <code>['flattop', 'hamming', ('kaiser', 4.0), 'blackman', 'hann']</code>  , hann window is used at test and validation time</li>\n<li>masking out non labeled chunks of the audio with a 10% chance</li>\n<li>one of spectrogram <a href=\"https://arxiv.org/abs/2002.12047\" target=\"_blank\">FMix</a> and audio based mixup with the max of labels instead of using the blend from the beta parameter</li>\n<li><a href=\"http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf\" target=\"_blank\">spec mix </a>: <br>\nonly one strip was used for each axis, for the horizontal axis when the chosen frequency range to mask out completely covers a specific species minimum f_min and maximum f_max , that species label will be dropped. Specmix is also using the max of labels instead of using the blend from the beta parameter. The code below shows how I obtain the function that can output frequency axis spectrogram positions from frequency</li>\n</ul>\n<pre><code>def get_mel_scaled_hz_to_y_axis_func(fmin=0, fmax=16000, n_mels=128):\n    hz_points = librosa.core.mel_frequencies(n_mels=n_mels, fmin=fmin, fmax=fmax)\n    hz_to_y_axis = interp1d(hz_points, np.arange(n_mels)[::-1])  \n    # reversed because first index is at the top left in an image array\n    return hz_to_y_axis\n</code></pre>\n<ul>\n<li>bandpass noise </li>\n<li>Water from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Engine and Motor Sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Honk, Traffic and Horn sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Speech sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n<li>Bark sounds from <a href=\"https://zenodo.org/record/4060432#.YC6IJ3VKjCI\" target=\"_blank\">Freesound50k</a> removing samples that have license to prevent derivative work</li>\n</ul>\n<p>checkout recording_id b8d1e4865 to find dogs barking and some human speech :D</p>\n<h6>Architectures used</h6>\n<p>No SED just plain classifier models with GEM pooling for CNN based models</p>\n<ul>\n<li>Efficientnet-b7</li>\n<li>Efficientnet-b8</li>\n<li>HRNet w64</li>\n<li>deitbase224</li>\n<li>vit_large_patch16_224</li>\n<li>ecaresnet50</li>\n<li>2x resnest50 from <a href=\"https://www.kaggle.com/meaninglesslives\" target=\"_blank\">https://www.kaggle.com/meaninglesslives</a>, checkout his writeup in a minimal notebook <a href=\"https://www.kaggle.com/meaninglesslives/rfcx-minimal\" target=\"_blank\">HERE</a>!</li>\n</ul>\n<h6>Loss</h6>\n<p>The main loss strategy used for the final submission was using different loss function for samples which I am confident is complete in labels and samples which I am not confident is complete in labels. BCE was used for non-noisy/confident samples and a modified Lsoft loss was used on the noisy/non-confident. Lsoft loss was modified to be applied only to nonpositive samples, as I was confident in my manual labels. It looks like this </p>\n<pre><code>def l_soft_on_negative_samples(y_pred, y_true, beta, eps = 1e-7):\n    y_pred = torch.clamp(y_pred, eps, 1.0)\n\n    # (1) dynamically update the targets based on the current state of the model:\n    # bootstrapped target tensor\n    # use predicted class proba directly to generate regression targets\n    with torch.no_grad():\n        negative_indexes = (y_true == 0).nonzero().squeeze(1)\n        y_true_update = y_true\n        y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] = (\n            y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] * beta + \n            (1 - beta) * y_pred[negative_indexes[:, 0], negative_indexes[:, 1]]\n        )\n\n    # (2) compute loss as always\n    loss = F.binary_cross_entropy(y_pred, y_true_update)\n    return loss\n</code></pre>\n<p>This was inspired by the first placed winner in the Freesound competition <a href=\"https://github.com/lRomul/argus-freesound\" target=\"_blank\">https://github.com/lRomul/argus-freesound</a> but I noticed that it doesn't make sense if it is used with mixup since audio will be mixed up anyways. So I also obtain the max of noisy binary labels so that noisy labels mixed with clean labels are considered to be noisy labels.</p>\n<h6>Pseudo Labels</h6>\n<p>I didn't get much boost from pseudo labels, maybe I did something wrong but nonetheless, it was used in some models. I used a 0.8 threshold for labels generated with 5-second windows and utilized the same window positions during training. Using raw predictions didn't help the model at all on lb.</p>\n<h6>Post processing</h6>\n<p>We set the species 3 labels to be 1 with a 0.95 threshold and it boosted the score slightly</p>\n<h6>Other stuff</h6>\n<ul>\n<li>Early stopping of 20 epochs with a minimum learning rate of 9e-6 to start counting these 20 epochs</li>\n<li>Reduce learning rate on Plateau with a factor of 0.6 and start with a few warmup epochs, when LR is reduced the best model weights was loaded back again</li>\n<li>Catalyst was used</li>\n</ul>\n<h2>Things that failed</h2>\n<ul>\n<li>using models without pre-trained weights</li>\n<li>timeshift</li>\n<li>using species from the Cornell Competition that are confused with species in this competition as a distractor noise, for example, moudov is similar to species 15, reevir is similar to species 11, rebwoo is similar to species 6, bkbwar is similar to species 7, cacwre is similar to species 19, every is similar to species 17 and nrwswa is similar to species 20</li>\n<li>using plane sounds from Freesound50k data</li>\n<li>using PCEN, deltas or CQT</li>\n<li><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">Random Power</a></li>\n<li>TTA with different window types</li>\n<li>Manifold mixup with resnest50</li>\n<li>using trainable Switchnorm as an initial layer replacing normal standardization</li>\n<li>using trainable Examplar norm as an initial layer replacing normal standardization</li>\n<li>Context Gating</li>\n<li>split audio intro three equal-length chunks and concat as 3 layer image</li>\n<li>lsep and <a href=\"https://arxiv.org/abs/2009.14119\" target=\"_blank\">Assymetric loss </a></li>\n<li>using rain sounds from Freesound50k data</li>\n<li>using a fixed validation mask similar to how I used random training mask</li>\n<li>use <a href=\"https://vsitzmann.github.io/siren/\" target=\"_blank\">SIREN</a> layer</li>\n<li>Tried to separate some confusing patterns as separate manual labels but didn't get the chance to test them</li>\n</ul>\n<p>Hopefully, I didn't miss anything. Oh, we were holding off submitting a good model until <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> came along :)</p>",
      "rawMarkdown": "# 3rd Place Solution\n\n**TLDR**\n\nOur solution is a mean blend of 8 models trained on True positive labels of all recording ids in  train_tp.csv (given + hand-labeled labels) also from some recording ids in train_fp.csv (hand-labeled labels) with heavy augmentations. Additionally, some models are also trained on pseudo labels and a hand-labeled external dataset. We also post-processed the blended results by thresholding species 3.\n\nFrom 308 submissions it is obvious that we have tested a lot of techniques and I will share more detailed information by category below.\n\n## Data Preparation\nI couldn't get a proper validation framework setup after trying out many techniques and decided at one point to start digging into the data and figured out that there are many unlabeled samples both in and out of the range of t_min and t_max labels given in train_tp.csv. In https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735 it was mentioned that using hand-labeled species was allowed, so I started labeling the data manually and after labeling 100 recording ids I could already get a > 0.9 public lb score and pretty consistent local CV scores that somewhat correlates with the public lb. Naturally, I continued to label the entire train_tp.csv seeing that it has only around 1.3k recording ids.  Further labeling of train_fp.csv helped the score but only minimally so I stopped at one point. As I grew more familiar with the data I could label 300 recording ids in a day :), referring to pseudo labels helped a lot too. I also went through the train_tp.csv a few more rounds to make sure I have quality data. I used both spectrograms and listening strategy to analyze and label the data, some species are easy to spot with spectrograms and some are easier to spot by listening, and in some cases, both listening and visual inspection of the spectrograms can act as a multi verification technique to get more quality labels, especially when birds/frogs are very distant away from the recorder or there are strong noises like waterfall sounds. By labeling and analyzing the data I also figured out the kinds of sounds/noises that would appear and inspired me to try out a few augmentation methods which I will share below. Along with true positive labels, I also added noisy and non-noisy labels based on my confidence in the completeness of labels in a specific recording id. I am not a perfect labeler so I wanted to handle complete and non-complete labeled recording ids differently, which I will share below too.\n\nI also removed some labels from train_tp.csv as I found some true positives suspicious, I didn't test not removing the labels before so not sure how much this helped.\n\nAdditionally, after finding out the paper from the organizers I searched for suitably licensed datasets with those species and found one dataset with species in this competition with a proper license. But there weren't any labels so I labeled it manually too with the same format as train_tp.csv. https://datadryad.org/stash/dataset/doi:10.5061/dryad.c0g2t . I reuploaded the dataset [HERE](https://www.kaggle.com/dicksonchin93/eleutherodactylus-frogs) with my manual labels.\n\nI uploaded the extra labels as a dataset https://www.kaggle.com/dicksonchin93/extra-labels-for-rcfx-competition-data, feel free to use it and see if you can get a better single model score! mine was 0.970 on public lb\n \n## Modeling / Data Pre-processing \n\nI used Mel Spectrograms with the following parameters: 32kHz sampling rate, a hop size of 716, a window size of 1366, and 224 or 128 Number of Mels. Tried a bunch of methods but plainly using 3 layers of standardized Mel Spectrograms works the best. The image dimensions were (num_mel_bins, 750). \n\nUsing train_tp.csv to create folds will potentially leak some training data into your validation data so I treated the problem as a multilabel target and used [iterative-stratification](https://github.com/trent-b/iterative-stratification) to stratify the data into 5 partitions using unique recording ids and its multilabel targets. I had two different 5 fold partitions using different versions of the multi labels and used a mix of both in the final submission. \n\nI used multiple different audio duration during the competition and at different stages of the competition, the best duration varied in my implementation but in the end, I used 5 seconds of audio for training and prediction as the LWLWRAP score was better on both public lb and local validation. \n\nThe 5-second audio was randomly sampled during training and in prediction time a 5-second sliding window was used with overlap and the max of predictions was used. How the 5-second audio is randomly sampled is considered to be an augmentation method in my opinion and so I will explain it in the heavy augmentations category below\n\n###### Augmentations\n\n- Random 5-second audio samples: \na starting point was chosen randomly on values between reference t_mins and t_maxes obtained from \n\n```\ndef get_ref_tmin_tmax_and_species_ids(\n    self, all_tp_events, label_column_key=\"species_id\"\n):\n        all_tp_events[\"t_min_ref\"] = all_tp_events[\"t_min\"].apply(\n           lambda x: max(x - (self.period / 2.0), 0)\n        )\n        def get_tmax_ref(row, period=self.period):\n            tmin_x = row[\"t_min\"]\n            tmax_x = row[\"t_max\"]\n            tmax_ref = tmax_x - (period / 4.0)\n            if tmax_ref < tmin_x:\n                tmax_ref = (tmax_x - tmin_x) / 2.0 + tmin_x\n            return tmax_ref\n        all_tp_events[\"t_max_ref\"] = all_tp_events[\n            [\"t_max\", \"t_min\"]\n        ].apply(get_tmax_ref, axis=1)\n        t_min_maxes = all_tp_events[\n            [\"t_min_ref\", \"t_max_ref\"]\n        ].values.tolist()\n        species_ids = all_tp_events[label_column_key].values.tolist()\n        return t_min_maxes, species_ids\n\n```\nLabels were also assigned based on the chosen starting time and ending time with t_min and t_max labels.\n- audio based pink noise\n- audio based white noise\n- reverberation\n- time stretch\n- use one of 16kHz or 48kHz sample rate data and resample it to 32kHz sample rate using randomly chosen resampling methods `['kaiser_best', 'kaiser_fast', 'fft', 'polyphase']`\n- use different window types to compute spectrograms at train time `['flattop', 'hamming', ('kaiser', 4.0), 'blackman', 'hann'] `  , hann window is used at test and validation time\n- masking out non labeled chunks of the audio with a 10% chance\n- one of spectrogram [FMix](https://arxiv.org/abs/2002.12047) and audio based mixup with the max of labels instead of using the blend from the beta parameter\n- [spec mix ](http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf): \nonly one strip was used for each axis, for the horizontal axis when the chosen frequency range to mask out completely covers a specific species minimum f_min and maximum f_max , that species label will be dropped. Specmix is also using the max of labels instead of using the blend from the beta parameter. The code below shows how I obtain the function that can output frequency axis spectrogram positions from frequency\n```\ndef get_mel_scaled_hz_to_y_axis_func(fmin=0, fmax=16000, n_mels=128):\n    hz_points = librosa.core.mel_frequencies(n_mels=n_mels, fmin=fmin, fmax=fmax)\n    hz_to_y_axis = interp1d(hz_points, np.arange(n_mels)[::-1])  \n    # reversed because first index is at the top left in an image array\n    return hz_to_y_axis\n```\n- bandpass noise \n- Water from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Engine and Motor Sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Honk, Traffic and Horn sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Speech sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Bark sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n\n\ncheckout recording_id b8d1e4865 to find dogs barking and some human speech :D\n\n###### Architectures used\nNo SED just plain classifier models with GEM pooling for CNN based models\n- Efficientnet-b7\n- Efficientnet-b8\n- HRNet w64\n- deitbase224\n- vit_large_patch16_224\n- ecaresnet50\n- 2x resnest50 from https://www.kaggle.com/meaninglesslives, checkout his writeup in a minimal notebook [HERE](https://www.kaggle.com/meaninglesslives/rfcx-minimal)!\n\n###### Loss\n\nThe main loss strategy used for the final submission was using different loss function for samples which I am confident is complete in labels and samples which I am not confident is complete in labels. BCE was used for non-noisy/confident samples and a modified Lsoft loss was used on the noisy/non-confident. Lsoft loss was modified to be applied only to nonpositive samples, as I was confident in my manual labels. It looks like this \n```\ndef l_soft_on_negative_samples(y_pred, y_true, beta, eps = 1e-7):\n    y_pred = torch.clamp(y_pred, eps, 1.0)\n\n    # (1) dynamically update the targets based on the current state of the model:\n    # bootstrapped target tensor\n    # use predicted class proba directly to generate regression targets\n    with torch.no_grad():\n        negative_indexes = (y_true == 0).nonzero().squeeze(1)\n        y_true_update = y_true\n        y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] = (\n            y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] * beta + \n            (1 - beta) * y_pred[negative_indexes[:, 0], negative_indexes[:, 1]]\n        )\n        \n    # (2) compute loss as always\n    loss = F.binary_cross_entropy(y_pred, y_true_update)\n    return loss\n```\nThis was inspired by the first placed winner in the Freesound competition https://github.com/lRomul/argus-freesound but I noticed that it doesn't make sense if it is used with mixup since audio will be mixed up anyways. So I also obtain the max of noisy binary labels so that noisy labels mixed with clean labels are considered to be noisy labels.\n\n###### Pseudo Labels\nI didn't get much boost from pseudo labels, maybe I did something wrong but nonetheless, it was used in some models. I used a 0.8 threshold for labels generated with 5-second windows and utilized the same window positions during training. Using raw predictions didn't help the model at all on lb.\n\n###### Post processing\nWe set the species 3 labels to be 1 with a 0.95 threshold and it boosted the score slightly\n\n\n###### Other stuff\n\n- Early stopping of 20 epochs with a minimum learning rate of 9e-6 to start counting these 20 epochs\n- Reduce learning rate on Plateau with a factor of 0.6 and start with a few warmup epochs, when LR is reduced the best model weights was loaded back again\n- Catalyst was used\n\n## Things that failed\n\n- using models without pre-trained weights\n- timeshift\n- using species from the Cornell Competition that are confused with species in this competition as a distractor noise, for example, moudov is similar to species 15, reevir is similar to species 11, rebwoo is similar to species 6, bkbwar is similar to species 7, cacwre is similar to species 19, every is similar to species 17 and nrwswa is similar to species 20\n- using plane sounds from Freesound50k data\n- using PCEN, deltas or CQT\n- [Random Power](https://www.kaggle.com/c/birdsong-recognition/discussion/183269)\n- TTA with different window types\n- Manifold mixup with resnest50\n- using trainable Switchnorm as an initial layer replacing normal standardization\n- using trainable Examplar norm as an initial layer replacing normal standardization\n- Context Gating\n- split audio intro three equal-length chunks and concat as 3 layer image\n- lsep and [Assymetric loss ](https://arxiv.org/abs/2009.14119)\n- using rain sounds from Freesound50k data\n- using a fixed validation mask similar to how I used random training mask\n- use [SIREN](https://vsitzmann.github.io/siren/) layer\n- Tried to separate some confusing patterns as separate manual labels but didn't get the chance to test them\n\n\nHopefully, I didn't miss anything. Oh, we were holding off submitting a good model until @cpmpml came along :)",
      "votes": null
    },
    {
      "id": "1209241",
      "postDate": "02/18/2021 19:12:04",
      "content": "<p>Thank you very much for sharing your approach, but it seems that hand-labeling was not allowed</p>",
      "rawMarkdown": "Thank you very much for sharing your approach, but it seems that hand-labeling was not allowed",
      "votes": null
    },
    {
      "id": "1209249",
      "postDate": "02/18/2021 19:14:26",
      "content": "<p>nope, only hand labeling test set is not allowed <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735</a></p>",
      "rawMarkdown": "nope, only hand labeling test set is not allowed https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735",
      "votes": null
    },
    {
      "id": "1209251",
      "postDate": "02/18/2021 19:15:23",
      "content": "<p>oh, i see, thanks!</p>",
      "rawMarkdown": "oh, i see, thanks!",
      "votes": null
    },
    {
      "id": "1209257",
      "postDate": "02/18/2021 19:22:41",
      "content": "<p>Hand labeling is allowed, as shown in link above.</p>\n<p>But 7C in competition rules - \"However, you will ensure the External Data is publicly available to use by all participants of the Competition for purposes of the competition at no cost to the other participants.\"</p>\n<p>But then there are no more external data sharing thread. Possibly outdated rule? Could hand labels be even considered as external data?</p>",
      "rawMarkdown": "Hand labeling is allowed, as shown in link above.\n\nBut 7C in competition rules - \"However, you will ensure the External Data is publicly available to use by all participants of the Competition for purposes of the competition at no cost to the other participants.\"\n\nBut then there are no more external data sharing thread. Possibly outdated rule? Could hand labels be even considered as external data?",
      "votes": null
    },
    {
      "id": "1209267",
      "postDate": "02/18/2021 19:35:19",
      "content": "<p>Data only has to be accessible to everybody, you just don't have to tell everybody which datasets you're using anymore. So their approach is totally fine I believe :)</p>",
      "rawMarkdown": "Data only has to be accessible to everybody, you just don't have to tell everybody which datasets you're using anymore. So their approach is totally fine I believe :)",
      "votes": null
    },
    {
      "id": "1209386",
      "postDate": "02/18/2021 21:29:14",
      "content": "<p>Wow, lots of work here.  I'll ned to read carefully again. Congrats on the ending.</p>\n<p>I share some fails, esp these ones:</p>\n<ul>\n<li>using PCEN, deltas or CQT</li>\n<li>random power</li>\n</ul>",
      "rawMarkdown": "Wow, lots of work here.  I'll ned to read carefully again. Congrats on the ending.\n\nI share some fails, esp these ones:\n- using PCEN, deltas or CQT\n- random power",
      "votes": null
    },
    {
      "id": "1209400",
      "postDate": "02/18/2021 21:37:37",
      "content": "<p>Thanks! and yeah, imagine the stress when you did so much work and someone came along and said they got top 10 in one submission 😨</p>",
      "rawMarkdown": "Thanks! and yeah, imagine the stress when you did so much work and someone came along and said they got top 10 in one submission 😨",
      "votes": null
    },
    {
      "id": "1209466",
      "postDate": "02/18/2021 23:10:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/dicksonchin93\" target=\"_blank\">@dicksonchin93</a> and team. Thanks for sharing details solution and code, lots of work! </p>",
      "rawMarkdown": "Congrats @dicksonchin93 and team. Thanks for sharing details solution and code, lots of work!",
      "votes": null
    },
    {
      "id": "1209480",
      "postDate": "02/18/2021 23:24:01",
      "content": "<p>Congrats. I'm glad your hard work is success.</p>",
      "rawMarkdown": "Congrats. I'm glad your hard work is success.",
      "votes": null
    },
    {
      "id": "1209484",
      "postDate": "02/18/2021 23:26:08",
      "content": "<p>Thanks Shinmura-san :)</p>",
      "rawMarkdown": "Thanks Shinmura-san :)",
      "votes": null
    },
    {
      "id": "1209653",
      "postDate": "02/19/2021 01:55:59",
      "content": "<p>May I ask for your training setup for \"vit_large_patch16_224\"? </p>\n<p>I tried 'vit_base_patch16_224' and found it difficult to converge (it barely reached 0.75 at validation with SGD and won't converge at all with adam). The batch size is 32, is this too small maybe?</p>\n<p>Thanks and congrats on the result!</p>",
      "rawMarkdown": "May I ask for your training setup for \"vit_large_patch16_224\"? \n\nI tried 'vit_base_patch16_224' and found it difficult to converge (it barely reached 0.75 at validation with SGD and won't converge at all with adam). The batch size is 32, is this too small maybe?\n\nThanks and congrats on the result!",
      "votes": null
    },
    {
      "id": "1209661",
      "postDate": "02/19/2021 02:06:24",
      "content": "<p>Thanks! I couldn't run vit_large_patch16_224 with a large batch size because of all my techniques, I only used batch size 10 but I don't think batch size is the problem, I couldn't get it to converge initially too, and decided to try a lower learning rate (0.0001 instead of 0.001) and it could slowly converge nicely. </p>",
      "rawMarkdown": "Thanks! I couldn't run vit_large_patch16_224 with a large batch size because of all my techniques, I only used batch size 10 but I don't think batch size is the problem, I couldn't get it to converge initially too, and decided to try a lower learning rate (0.0001 instead of 0.001) and it could slowly converge nicely.",
      "votes": null
    },
    {
      "id": "1210589",
      "postDate": "02/19/2021 14:38:50",
      "content": "<p>Congratulations, and nice work 👍 </p>",
      "rawMarkdown": "Congratulations, and nice work 👍",
      "votes": null
    },
    {
      "id": "1210736",
      "postDate": "02/19/2021 16:35:15",
      "content": "<p>forgot to mention that the initial few epochs with learning rate warmup helped too</p>",
      "rawMarkdown": "forgot to mention that the initial few epochs with learning rate warmup helped too",
      "votes": null
    },
    {
      "id": "1212414",
      "postDate": "02/21/2021 07:54:04",
      "content": "<p>Where can I find the organizer's paper?<br>\nI have not been able to find it.</p>",
      "rawMarkdown": "Where can I find the organizer's paper?\nI have not been able to find it.",
      "votes": null
    },
    {
      "id": "1212554",
      "postDate": "02/21/2021 10:31:53",
      "content": "<blockquote>\n  <p>Where can I find the organizer's paper?</p>\n</blockquote>\n<p>Check <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735</a></p>",
      "rawMarkdown": "> Where can I find the organizer's paper?\n\nCheck https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735",
      "votes": null
    },
    {
      "id": "1212560",
      "postDate": "02/21/2021 10:44:48",
      "content": "<p>Oh interesting! Thanks for the info</p>",
      "rawMarkdown": "Oh interesting! Thanks for the info",
      "votes": null
    },
    {
      "id": "1212587",
      "postDate": "02/21/2021 11:12:55",
      "content": "<p>Thank you!!</p>",
      "rawMarkdown": "Thank you!!",
      "votes": null
    },
    {
      "id": "1212816",
      "postDate": "02/21/2021 16:22:27",
      "content": "<p>Wow congratulations, too much hard work here in this solution. Thankfully, it ended up with the success 👍</p>",
      "rawMarkdown": "Wow congratulations, too much hard work here in this solution. Thankfully, it ended up with the success 👍",
      "votes": null
    },
    {
      "id": "1219889",
      "postDate": "02/27/2021 10:18:17",
      "content": "<p>Congrats, and good to see your hard work gets a great result! </p>\n<blockquote>\n  <p>As I grew more familiar with the data I could label 300 recording ids in a day :)</p>\n</blockquote>\n<p>Could you share the tool you used to manually label the data? I think you must have a well-designed tool so you can label such fast.</p>",
      "rawMarkdown": "Congrats, and good to see your hard work gets a great result! \n\n> As I grew more familiar with the data I could label 300 recording ids in a day :)\n\nCould you share the tool you used to manually label the data? I think you must have a well-designed tool so you can label such fast.",
      "votes": null
    },
    {
      "id": "1220094",
      "postDate": "02/27/2021 14:40:59",
      "content": "<p>Thanks! no custom tool was used, I just used a notebook to visualize the 60 seconds spectrogram(with an adequate amount of xticks and yticks along with original tp and fp bounding box annotations) and to play the audio file while referring to an unaggregated visualization of the 5-second ensembled model pseudo labels</p>",
      "rawMarkdown": "Thanks! no custom tool was used, I just used a notebook to visualize the 60 seconds spectrogram(with an adequate amount of xticks and yticks along with original tp and fp bounding box annotations) and to play the audio file while referring to an unaggregated visualization of the 5-second ensembled model pseudo labels",
      "votes": null
    },
    {
      "id": "1220200",
      "postDate": "02/27/2021 17:54:07",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1209241,
      "author_name": "loompa",
      "author_url": "",
      "post_date": "02/18/2021 19:12:04",
      "content": "<p>Thank you very much for sharing your approach, but it seems that hand-labeling was not allowed</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209249,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/18/2021 19:14:26",
          "content": "<p>nope, only hand labeling test set is not allowed <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209251,
          "author_name": "loompa",
          "author_url": "",
          "post_date": "02/18/2021 19:15:23",
          "content": "<p>oh, i see, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209257,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "02/18/2021 19:22:41",
          "content": "<p>Hand labeling is allowed, as shown in link above.</p>\n<p>But 7C in competition rules - \"However, you will ensure the External Data is publicly available to use by all participants of the Competition for purposes of the competition at no cost to the other participants.\"</p>\n<p>But then there are no more external data sharing thread. Possibly outdated rule? Could hand labels be even considered as external data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209267,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/18/2021 19:35:19",
          "content": "<p>Data only has to be accessible to everybody, you just don't have to tell everybody which datasets you're using anymore. So their approach is totally fine I believe :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209386,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 21:29:14",
      "content": "<p>Wow, lots of work here.  I'll ned to read carefully again. Congrats on the ending.</p>\n<p>I share some fails, esp these ones:</p>\n<ul>\n<li>using PCEN, deltas or CQT</li>\n<li>random power</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1209400,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/18/2021 21:37:37",
          "content": "<p>Thanks! and yeah, imagine the stress when you did so much work and someone came along and said they got top 10 in one submission 😨</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209466,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 23:10:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/dicksonchin93\" target=\"_blank\">@dicksonchin93</a> and team. Thanks for sharing details solution and code, lots of work! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209480,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "02/18/2021 23:24:01",
      "content": "<p>Congrats. I'm glad your hard work is success.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209484,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/18/2021 23:26:08",
          "content": "<p>Thanks Shinmura-san :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209653,
      "author_name": "dannywu375",
      "author_url": "",
      "post_date": "02/19/2021 01:55:59",
      "content": "<p>May I ask for your training setup for \"vit_large_patch16_224\"? </p>\n<p>I tried 'vit_base_patch16_224' and found it difficult to converge (it barely reached 0.75 at validation with SGD and won't converge at all with adam). The batch size is 32, is this too small maybe?</p>\n<p>Thanks and congrats on the result!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209661,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/19/2021 02:06:24",
          "content": "<p>Thanks! I couldn't run vit_large_patch16_224 with a large batch size because of all my techniques, I only used batch size 10 but I don't think batch size is the problem, I couldn't get it to converge initially too, and decided to try a lower learning rate (0.0001 instead of 0.001) and it could slowly converge nicely. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210736,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/19/2021 16:35:15",
          "content": "<p>forgot to mention that the initial few epochs with learning rate warmup helped too</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212560,
          "author_name": "dannywu375",
          "author_url": "",
          "post_date": "02/21/2021 10:44:48",
          "content": "<p>Oh interesting! Thanks for the info</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210589,
      "author_name": "milobele",
      "author_url": "",
      "post_date": "02/19/2021 14:38:50",
      "content": "<p>Congratulations, and nice work 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1212414,
      "author_name": "yururoi",
      "author_url": "",
      "post_date": "02/21/2021 07:54:04",
      "content": "<p>Where can I find the organizer's paper?<br>\nI have not been able to find it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212554,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "02/21/2021 10:31:53",
          "content": "<blockquote>\n  <p>Where can I find the organizer's paper?</p>\n</blockquote>\n<p>Check <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212587,
          "author_name": "yururoi",
          "author_url": "",
          "post_date": "02/21/2021 11:12:55",
          "content": "<p>Thank you!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212816,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/21/2021 16:22:27",
      "content": "<p>Wow congratulations, too much hard work here in this solution. Thankfully, it ended up with the success 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1220200,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/27/2021 17:54:07",
          "content": "<p>thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1219889,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "02/27/2021 10:18:17",
      "content": "<p>Congrats, and good to see your hard work gets a great result! </p>\n<blockquote>\n  <p>As I grew more familiar with the data I could label 300 recording ids in a day :)</p>\n</blockquote>\n<p>Could you share the tool you used to manually label the data? I think you must have a well-designed tool so you can label such fast.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1220094,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/27/2021 14:40:59",
          "content": "<p>Thanks! no custom tool was used, I just used a notebook to visualize the 60 seconds spectrogram(with an adequate amount of xticks and yticks along with original tp and fp bounding box annotations) and to play the audio file while referring to an unaggregated visualization of the 5-second ensembled model pseudo labels</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1209066": "# 3rd Place Solution\n\n**TLDR**\n\nOur solution is a mean blend of 8 models trained on True positive labels of all recording ids in  train_tp.csv (given + hand-labeled labels) also from some recording ids in train_fp.csv (hand-labeled labels) with heavy augmentations. Additionally, some models are also trained on pseudo labels and a hand-labeled external dataset. We also post-processed the blended results by thresholding species 3.\n\nFrom 308 submissions it is obvious that we have tested a lot of techniques and I will share more detailed information by category below.\n\n## Data Preparation\nI couldn't get a proper validation framework setup after trying out many techniques and decided at one point to start digging into the data and figured out that there are many unlabeled samples both in and out of the range of t_min and t_max labels given in train_tp.csv. In https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735 it was mentioned that using hand-labeled species was allowed, so I started labeling the data manually and after labeling 100 recording ids I could already get a > 0.9 public lb score and pretty consistent local CV scores that somewhat correlates with the public lb. Naturally, I continued to label the entire train_tp.csv seeing that it has only around 1.3k recording ids.  Further labeling of train_fp.csv helped the score but only minimally so I stopped at one point. As I grew more familiar with the data I could label 300 recording ids in a day :), referring to pseudo labels helped a lot too. I also went through the train_tp.csv a few more rounds to make sure I have quality data. I used both spectrograms and listening strategy to analyze and label the data, some species are easy to spot with spectrograms and some are easier to spot by listening, and in some cases, both listening and visual inspection of the spectrograms can act as a multi verification technique to get more quality labels, especially when birds/frogs are very distant away from the recorder or there are strong noises like waterfall sounds. By labeling and analyzing the data I also figured out the kinds of sounds/noises that would appear and inspired me to try out a few augmentation methods which I will share below. Along with true positive labels, I also added noisy and non-noisy labels based on my confidence in the completeness of labels in a specific recording id. I am not a perfect labeler so I wanted to handle complete and non-complete labeled recording ids differently, which I will share below too.\n\nI also removed some labels from train_tp.csv as I found some true positives suspicious, I didn't test not removing the labels before so not sure how much this helped.\n\nAdditionally, after finding out the paper from the organizers I searched for suitably licensed datasets with those species and found one dataset with species in this competition with a proper license. But there weren't any labels so I labeled it manually too with the same format as train_tp.csv. https://datadryad.org/stash/dataset/doi:10.5061/dryad.c0g2t . I reuploaded the dataset [HERE](https://www.kaggle.com/dicksonchin93/eleutherodactylus-frogs) with my manual labels.\n\nI uploaded the extra labels as a dataset https://www.kaggle.com/dicksonchin93/extra-labels-for-rcfx-competition-data, feel free to use it and see if you can get a better single model score! mine was 0.970 on public lb\n \n## Modeling / Data Pre-processing \n\nI used Mel Spectrograms with the following parameters: 32kHz sampling rate, a hop size of 716, a window size of 1366, and 224 or 128 Number of Mels. Tried a bunch of methods but plainly using 3 layers of standardized Mel Spectrograms works the best. The image dimensions were (num_mel_bins, 750). \n\nUsing train_tp.csv to create folds will potentially leak some training data into your validation data so I treated the problem as a multilabel target and used [iterative-stratification](https://github.com/trent-b/iterative-stratification) to stratify the data into 5 partitions using unique recording ids and its multilabel targets. I had two different 5 fold partitions using different versions of the multi labels and used a mix of both in the final submission. \n\nI used multiple different audio duration during the competition and at different stages of the competition, the best duration varied in my implementation but in the end, I used 5 seconds of audio for training and prediction as the LWLWRAP score was better on both public lb and local validation. \n\nThe 5-second audio was randomly sampled during training and in prediction time a 5-second sliding window was used with overlap and the max of predictions was used. How the 5-second audio is randomly sampled is considered to be an augmentation method in my opinion and so I will explain it in the heavy augmentations category below\n\n###### Augmentations\n\n- Random 5-second audio samples: \na starting point was chosen randomly on values between reference t_mins and t_maxes obtained from \n\n```\ndef get_ref_tmin_tmax_and_species_ids(\n    self, all_tp_events, label_column_key=\"species_id\"\n):\n        all_tp_events[\"t_min_ref\"] = all_tp_events[\"t_min\"].apply(\n           lambda x: max(x - (self.period / 2.0), 0)\n        )\n        def get_tmax_ref(row, period=self.period):\n            tmin_x = row[\"t_min\"]\n            tmax_x = row[\"t_max\"]\n            tmax_ref = tmax_x - (period / 4.0)\n            if tmax_ref < tmin_x:\n                tmax_ref = (tmax_x - tmin_x) / 2.0 + tmin_x\n            return tmax_ref\n        all_tp_events[\"t_max_ref\"] = all_tp_events[\n            [\"t_max\", \"t_min\"]\n        ].apply(get_tmax_ref, axis=1)\n        t_min_maxes = all_tp_events[\n            [\"t_min_ref\", \"t_max_ref\"]\n        ].values.tolist()\n        species_ids = all_tp_events[label_column_key].values.tolist()\n        return t_min_maxes, species_ids\n\n```\nLabels were also assigned based on the chosen starting time and ending time with t_min and t_max labels.\n- audio based pink noise\n- audio based white noise\n- reverberation\n- time stretch\n- use one of 16kHz or 48kHz sample rate data and resample it to 32kHz sample rate using randomly chosen resampling methods `['kaiser_best', 'kaiser_fast', 'fft', 'polyphase']`\n- use different window types to compute spectrograms at train time `['flattop', 'hamming', ('kaiser', 4.0), 'blackman', 'hann'] `  , hann window is used at test and validation time\n- masking out non labeled chunks of the audio with a 10% chance\n- one of spectrogram [FMix](https://arxiv.org/abs/2002.12047) and audio based mixup with the max of labels instead of using the blend from the beta parameter\n- [spec mix ](http://dcase.community/documents/challenge2019/technical_reports/DCASE2019_Bouteillon_27_t2.pdf): \nonly one strip was used for each axis, for the horizontal axis when the chosen frequency range to mask out completely covers a specific species minimum f_min and maximum f_max , that species label will be dropped. Specmix is also using the max of labels instead of using the blend from the beta parameter. The code below shows how I obtain the function that can output frequency axis spectrogram positions from frequency\n```\ndef get_mel_scaled_hz_to_y_axis_func(fmin=0, fmax=16000, n_mels=128):\n    hz_points = librosa.core.mel_frequencies(n_mels=n_mels, fmin=fmin, fmax=fmax)\n    hz_to_y_axis = interp1d(hz_points, np.arange(n_mels)[::-1])  \n    # reversed because first index is at the top left in an image array\n    return hz_to_y_axis\n```\n- bandpass noise \n- Water from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Engine and Motor Sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Honk, Traffic and Horn sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Speech sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n- Bark sounds from [Freesound50k](https://zenodo.org/record/4060432#.YC6IJ3VKjCI) removing samples that have license to prevent derivative work\n\n\ncheckout recording_id b8d1e4865 to find dogs barking and some human speech :D\n\n###### Architectures used\nNo SED just plain classifier models with GEM pooling for CNN based models\n- Efficientnet-b7\n- Efficientnet-b8\n- HRNet w64\n- deitbase224\n- vit_large_patch16_224\n- ecaresnet50\n- 2x resnest50 from https://www.kaggle.com/meaninglesslives, checkout his writeup in a minimal notebook [HERE](https://www.kaggle.com/meaninglesslives/rfcx-minimal)!\n\n###### Loss\n\nThe main loss strategy used for the final submission was using different loss function for samples which I am confident is complete in labels and samples which I am not confident is complete in labels. BCE was used for non-noisy/confident samples and a modified Lsoft loss was used on the noisy/non-confident. Lsoft loss was modified to be applied only to nonpositive samples, as I was confident in my manual labels. It looks like this \n```\ndef l_soft_on_negative_samples(y_pred, y_true, beta, eps = 1e-7):\n    y_pred = torch.clamp(y_pred, eps, 1.0)\n\n    # (1) dynamically update the targets based on the current state of the model:\n    # bootstrapped target tensor\n    # use predicted class proba directly to generate regression targets\n    with torch.no_grad():\n        negative_indexes = (y_true == 0).nonzero().squeeze(1)\n        y_true_update = y_true\n        y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] = (\n            y_true_update[negative_indexes[:, 0], negative_indexes[:, 1]] * beta + \n            (1 - beta) * y_pred[negative_indexes[:, 0], negative_indexes[:, 1]]\n        )\n        \n    # (2) compute loss as always\n    loss = F.binary_cross_entropy(y_pred, y_true_update)\n    return loss\n```\nThis was inspired by the first placed winner in the Freesound competition https://github.com/lRomul/argus-freesound but I noticed that it doesn't make sense if it is used with mixup since audio will be mixed up anyways. So I also obtain the max of noisy binary labels so that noisy labels mixed with clean labels are considered to be noisy labels.\n\n###### Pseudo Labels\nI didn't get much boost from pseudo labels, maybe I did something wrong but nonetheless, it was used in some models. I used a 0.8 threshold for labels generated with 5-second windows and utilized the same window positions during training. Using raw predictions didn't help the model at all on lb.\n\n###### Post processing\nWe set the species 3 labels to be 1 with a 0.95 threshold and it boosted the score slightly\n\n\n###### Other stuff\n\n- Early stopping of 20 epochs with a minimum learning rate of 9e-6 to start counting these 20 epochs\n- Reduce learning rate on Plateau with a factor of 0.6 and start with a few warmup epochs, when LR is reduced the best model weights was loaded back again\n- Catalyst was used\n\n## Things that failed\n\n- using models without pre-trained weights\n- timeshift\n- using species from the Cornell Competition that are confused with species in this competition as a distractor noise, for example, moudov is similar to species 15, reevir is similar to species 11, rebwoo is similar to species 6, bkbwar is similar to species 7, cacwre is similar to species 19, every is similar to species 17 and nrwswa is similar to species 20\n- using plane sounds from Freesound50k data\n- using PCEN, deltas or CQT\n- [Random Power](https://www.kaggle.com/c/birdsong-recognition/discussion/183269)\n- TTA with different window types\n- Manifold mixup with resnest50\n- using trainable Switchnorm as an initial layer replacing normal standardization\n- using trainable Examplar norm as an initial layer replacing normal standardization\n- Context Gating\n- split audio intro three equal-length chunks and concat as 3 layer image\n- lsep and [Assymetric loss ](https://arxiv.org/abs/2009.14119)\n- using rain sounds from Freesound50k data\n- using a fixed validation mask similar to how I used random training mask\n- use [SIREN](https://vsitzmann.github.io/siren/) layer\n- Tried to separate some confusing patterns as separate manual labels but didn't get the chance to test them\n\n\nHopefully, I didn't miss anything. Oh, we were holding off submitting a good model until @cpmpml came along :)",
    "1209241": "Thank you very much for sharing your approach, but it seems that hand-labeling was not allowed",
    "1209249": "nope, only hand labeling test set is not allowed https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735",
    "1209251": "oh, i see, thanks!",
    "1209257": "Hand labeling is allowed, as shown in link above.\n\nBut 7C in competition rules - \"However, you will ensure the External Data is publicly available to use by all participants of the Competition for purposes of the competition at no cost to the other participants.\"\n\nBut then there are no more external data sharing thread. Possibly outdated rule? Could hand labels be even considered as external data?",
    "1209267": "Data only has to be accessible to everybody, you just don't have to tell everybody which datasets you're using anymore. So their approach is totally fine I believe :)",
    "1209386": "Wow, lots of work here.  I'll ned to read carefully again. Congrats on the ending.\n\nI share some fails, esp these ones:\n- using PCEN, deltas or CQT\n- random power",
    "1209400": "Thanks! and yeah, imagine the stress when you did so much work and someone came along and said they got top 10 in one submission 😨",
    "1209466": "Congrats @dicksonchin93 and team. Thanks for sharing details solution and code, lots of work!",
    "1209480": "Congrats. I'm glad your hard work is success.",
    "1209484": "Thanks Shinmura-san :)",
    "1209653": "May I ask for your training setup for \"vit_large_patch16_224\"? \n\nI tried 'vit_base_patch16_224' and found it difficult to converge (it barely reached 0.75 at validation with SGD and won't converge at all with adam). The batch size is 32, is this too small maybe?\n\nThanks and congrats on the result!",
    "1209661": "Thanks! I couldn't run vit_large_patch16_224 with a large batch size because of all my techniques, I only used batch size 10 but I don't think batch size is the problem, I couldn't get it to converge initially too, and decided to try a lower learning rate (0.0001 instead of 0.001) and it could slowly converge nicely.",
    "1210589": "Congratulations, and nice work 👍",
    "1210736": "forgot to mention that the initial few epochs with learning rate warmup helped too",
    "1212414": "Where can I find the organizer's paper?\nI have not been able to find it.",
    "1212554": "> Where can I find the organizer's paper?\n\nCheck https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197735",
    "1212560": "Oh interesting! Thanks for the info",
    "1212587": "Thank you!!",
    "1212816": "Wow congratulations, too much hard work here in this solution. Thankfully, it ended up with the success 👍",
    "1219889": "Congrats, and good to see your hard work gets a great result! \n\n> As I grew more familiar with the data I could label 300 recording ids in a day :)\n\nCould you share the tool you used to manually label the data? I think you must have a well-designed tool so you can label such fast.",
    "1220094": "Thanks! no custom tool was used, I just used a notebook to visualize the 60 seconds spectrogram(with an adequate amount of xticks and yticks along with original tp and fp bounding box annotations) and to play the audio file while referring to an unaggregated visualization of the 5-second ensembled model pseudo labels",
    "1220200": "thank you!"
  },
  "source": "meta"
}