{
  "id": 220306,
  "title": "27th place simple solution (0.932 public, 0.940 private) - dl_pipeline",
  "url": "/competitions/rfcx-species-audio-detection/writeups/miguel-pinto-27th-place-simple-solution-0-932-publ",
  "author_name": "",
  "post_date": "2021-02-18T22:05:03.070Z",
  "votes": 26,
  "comment_count": 6,
  "views": 0,
  "content": "<p><strong>In summary:</strong></p>\n<ul>\n<li>Best single model (0.925 public lb): densenet121 features + fastai head</li>\n<li>Loss function: cross entropy</li>\n<li>Sampling: 128x1024 crops around true positive samples</li>\n<li>Spectrogram parameters: n mels 128, hop length 640, sample rate 32000</li>\n<li>Augmentations: clipping distortion, pitch shift and mixup</li>\n<li>Inference: Predict on crops with a small width (e.g. 128x128 instead of 128x1024 used for training) and calculate the max probability for each of the 24 classes.</li>\n</ul>\n<p><strong>Introductory monologue</strong><br>\nFirst and foremost, this was an interesting competition and a good learning opportunity as it is often the case in Kaggle! One “problem” of this competition is that the test data was labeled with a different method and no samples of labeled test data were provided. This makes it difficult to get a sense of validation score and increases the danger of overfiting the public test results. In fact, I almost gave up on this competition when I realized this was the case. But eventually I decided to get back to it and work on a simple solution and on a python library – <strong>dl_pipeline</strong> ( <a href=\"https://github.com/mnpinto/dl_pipeline\" target=\"_blank\">https://github.com/mnpinto/dl_pipeline</a>) – that I will use as a general framework for future kaggle competitions in general. Initially, the idea for dl_pipeline was just to keep my code more organized and more reusable but I figured that maybe there’s also some value in sharing it. </p>\n<p><strong>Data preprocessing</strong><br>\nSave all wave files in npy files with sample rate of 32000 Hz to save time.</p>\n<pre><code>def audio2npy(file, path_save:Path, sample_rate=32_000):\n    path_save.mkdir(exist_ok=True, parents=True)\n    wave, _ = librosa.load(file, sr=sample_rate)\n    np.save(path_save/f'{file.stem}.npy', wave)\n</code></pre>\n<p>I didn't convert the audio to spectrograms right away since I still want the ability to use audio augmentation on the waveforms.</p>\n<p><strong>Augmentations and Spectrograms</strong></p>\n<ul>\n<li>First I create crops on the waveform including the true positive labels with a number of samples calculated so that the spectrogram will have a width of 1024. </li>\n</ul>\n<p><strong>Note:</strong> Cropping before applying the augmentations is much faster than the other way around.</p>\n<ul>\n<li>Then for the waveform augmentations I used the <strong>audiomentations</strong> library (<a href=\"https://github.com/iver56/audiomentations)\" target=\"_blank\">https://github.com/iver56/audiomentations)</a>. I ended up using just the following augmentations as based on public lb I didn't find that others were helping, although this would require a proper validation to take any conclusions. </li>\n</ul>\n<pre><code>def audio_augment(sample_rate, p=0.25):\n    return Pipeline([\n        ClippingDistortion(sample_rate, max_percentile_threshold=10, p=p),\n        PitchShift(sample_rate, min_semitones=-8, max_semitones=8, p=p),\n    ])\n</code></pre>\n<p><strong>Note:</strong> Some augmentations are much slower, for example, pitch shift and time stretch. When using those augmentations the probability of use makes a big difference in how long the training takes. </p>\n<ul>\n<li>Then I searched the fastest way to convert the audio to spectrograms in the GPU and I ended up using <strong>nnAudio</strong> (<a href=\"https://github.com/KinWaiCheuk/nnAudio)\" target=\"_blank\">https://github.com/KinWaiCheuk/nnAudio)</a>. Again, converting to spectrogram after the waveform is cropped is a nice gain in processing time.</li>\n</ul>\n<p><strong>Model</strong><br>\nI tried several models but the one that got me a better result in the public leaderboard was densenet121 and the second-best ResNeSt50. One particularity is that I use for all the models the fastai head with strong dropout.</p>\n<p>The fastai head (using <code>create_head(num_features*2, num_classes, ps=0.8)</code>).</p>\n<pre><code>(1): Sequential(\n      (0): AdaptiveConcatPool2d(\n        (ap): AdaptiveAvgPool2d(output_size=1)\n        (mp): AdaptiveMaxPool2d(output_size=1)\n      )\n      (1): Flatten(full=False)\n      (2): BatchNorm1d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (3): Dropout(p=0.4, inplace=False)\n      (4): Linear(in_features=2048, out_features=512, bias=False)\n      (5): ReLU(inplace=True)\n      (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (7): Dropout(p=0.8, inplace=False)\n      (8): Linear(in_features=512, out_features=24, bias=False)\n    )\n</code></pre>\n<p><strong>Training</strong><br>\nI guess code speaks more than words, particularly for those familiar with fastai:</p>\n<pre><code>bs = 32\nlearn = Learner(dls, model, loss_func=cross_entropy, metrics=[accuracy, lrap], cbs=cbs)\nlearn.to_fp16(clip=0.5);\nlearn.fit_one_cycle(30, 1e-3, wd=3e-2, div_final=10, div=10)\n</code></pre>\n<p>So in English, this is a one cycle learning rate schedule with <strong>30 epochs</strong> starting with lr=1e-4, increasing to 1e-3 and then decreasing back to 1e-4 following cosine anealing schedule. The <strong>loss function</strong> is the good old cross-entropy. Also, a <strong>weight decay</strong> of 3e-2 was used, a <strong>gradient clip</strong> of 0.5 and the train was done with <strong>mixed-precision</strong> so that my GTX 1080 can handle a <strong>batch size</strong> of 32 with 128x1024 image size.</p>\n<p>One training epoch takes about 1 minute on my GTX 1080,  I guess it's not bad considering that I'm doing waveform augmentations on CPU that even with p=0.25 take some time.</p>\n<p><strong>Inference</strong><br>\nThis is the fun part because it was almost by mistake that I realised that making inference with smaller tiles is way better. I presume that this is the case because I'm training with cross-entropy for a single label problem but the test data is labelled with multiple labels. By using smaller crops the predictions are more multilabel friendly. The reason I've been using cross-entropy instead of binary cross-entropy and sigmoid for the typical multilabel problem is that for me the convergence was much faster using the cross-entropy approach and with better results. Maybe I made a mistake somewhere I don't know, I didn't investigate it in much detail.</p>\n<ul>\n<li><p>Run predictions on crops of the spectrogram with a width of 64, 128 and 256 (remember training was done with 1024), calculate the max probability for each class for each case (64, 128, 256) and the average of the 3 cases. The average of the 3 gave me public lb 0.928 on my best single model that I describe above, compared to 0.925 for just the 128 width inference.</p></li>\n<li><p>The final solution with public lb 0.932 and private lb 0.940 is an ensemble of a few training interations with some modifications. (I will update this tomorrow with more information).</p></li>\n</ul>\n<p><strong>dl_pipeline</strong><br>\nAnd again the code for this solution is now public on this repo: <a href=\"https://github.com/mnpinto/dl_pipeline\" target=\"_blank\">https://github.com/mnpinto/dl_pipeline</a></p>\n<p>The following code should correspond to the best single model solution but I need to check if I didn't mess up anything when cleaning the code:</p>\n<pre><code>#!/bin/bash\narch='densenet121'\nmodel_name='model_0'\nsample_rate=32000\nn_mels=128\nhop_length=640\n\nfor fold in 0 1 2 3 4\ndo\n    echo \"Training $model for fold $fold\"\n    kaggle_rainforest2021 --fold $fold --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --bs 32 --head_ps 0.8 \\\n        --tile_width 1024 --mixup true &gt;&gt; log.train\ndone\n\nfor tw in 64 128 256\ndo\n    echo \"Generate predictions for $model with tile_width of $tw\"\n    kaggle_rainforest2021 --run_test true --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --tile_width $tw \\\n        --save_preds true &gt;&gt; log.predict\ndone\n</code></pre>\n<hr>\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": "1207597",
      "postDate": "02/18/2021 00:08:04",
      "content": "<p><strong>In summary:</strong></p>\n<ul>\n<li>Best single model (0.925 public lb): densenet121 features + fastai head</li>\n<li>Loss function: cross entropy</li>\n<li>Sampling: 128x1024 crops around true positive samples</li>\n<li>Spectrogram parameters: n mels 128, hop length 640, sample rate 32000</li>\n<li>Augmentations: clipping distortion, pitch shift and mixup</li>\n<li>Inference: Predict on crops with a small width (e.g. 128x128 instead of 128x1024 used for training) and calculate the max probability for each of the 24 classes.</li>\n</ul>\n<p><strong>Introductory monologue</strong><br>\nFirst and foremost, this was an interesting competition and a good learning opportunity as it is often the case in Kaggle! One “problem” of this competition is that the test data was labeled with a different method and no samples of labeled test data were provided. This makes it difficult to get a sense of validation score and increases the danger of overfiting the public test results. In fact, I almost gave up on this competition when I realized this was the case. But eventually I decided to get back to it and work on a simple solution and on a python library – <strong>dl_pipeline</strong> ( <a href=\"https://github.com/mnpinto/dl_pipeline\" target=\"_blank\">https://github.com/mnpinto/dl_pipeline</a>) – that I will use as a general framework for future kaggle competitions in general. Initially, the idea for dl_pipeline was just to keep my code more organized and more reusable but I figured that maybe there’s also some value in sharing it. </p>\n<p><strong>Data preprocessing</strong><br>\nSave all wave files in npy files with sample rate of 32000 Hz to save time.</p>\n<pre><code>def audio2npy(file, path_save:Path, sample_rate=32_000):\n    path_save.mkdir(exist_ok=True, parents=True)\n    wave, _ = librosa.load(file, sr=sample_rate)\n    np.save(path_save/f'{file.stem}.npy', wave)\n</code></pre>\n<p>I didn't convert the audio to spectrograms right away since I still want the ability to use audio augmentation on the waveforms.</p>\n<p><strong>Augmentations and Spectrograms</strong></p>\n<ul>\n<li>First I create crops on the waveform including the true positive labels with a number of samples calculated so that the spectrogram will have a width of 1024. </li>\n</ul>\n<p><strong>Note:</strong> Cropping before applying the augmentations is much faster than the other way around.</p>\n<ul>\n<li>Then for the waveform augmentations I used the <strong>audiomentations</strong> library (<a href=\"https://github.com/iver56/audiomentations)\" target=\"_blank\">https://github.com/iver56/audiomentations)</a>. I ended up using just the following augmentations as based on public lb I didn't find that others were helping, although this would require a proper validation to take any conclusions. </li>\n</ul>\n<pre><code>def audio_augment(sample_rate, p=0.25):\n    return Pipeline([\n        ClippingDistortion(sample_rate, max_percentile_threshold=10, p=p),\n        PitchShift(sample_rate, min_semitones=-8, max_semitones=8, p=p),\n    ])\n</code></pre>\n<p><strong>Note:</strong> Some augmentations are much slower, for example, pitch shift and time stretch. When using those augmentations the probability of use makes a big difference in how long the training takes. </p>\n<ul>\n<li>Then I searched the fastest way to convert the audio to spectrograms in the GPU and I ended up using <strong>nnAudio</strong> (<a href=\"https://github.com/KinWaiCheuk/nnAudio)\" target=\"_blank\">https://github.com/KinWaiCheuk/nnAudio)</a>. Again, converting to spectrogram after the waveform is cropped is a nice gain in processing time.</li>\n</ul>\n<p><strong>Model</strong><br>\nI tried several models but the one that got me a better result in the public leaderboard was densenet121 and the second-best ResNeSt50. One particularity is that I use for all the models the fastai head with strong dropout.</p>\n<p>The fastai head (using <code>create_head(num_features*2, num_classes, ps=0.8)</code>).</p>\n<pre><code>(1): Sequential(\n      (0): AdaptiveConcatPool2d(\n        (ap): AdaptiveAvgPool2d(output_size=1)\n        (mp): AdaptiveMaxPool2d(output_size=1)\n      )\n      (1): Flatten(full=False)\n      (2): BatchNorm1d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (3): Dropout(p=0.4, inplace=False)\n      (4): Linear(in_features=2048, out_features=512, bias=False)\n      (5): ReLU(inplace=True)\n      (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (7): Dropout(p=0.8, inplace=False)\n      (8): Linear(in_features=512, out_features=24, bias=False)\n    )\n</code></pre>\n<p><strong>Training</strong><br>\nI guess code speaks more than words, particularly for those familiar with fastai:</p>\n<pre><code>bs = 32\nlearn = Learner(dls, model, loss_func=cross_entropy, metrics=[accuracy, lrap], cbs=cbs)\nlearn.to_fp16(clip=0.5);\nlearn.fit_one_cycle(30, 1e-3, wd=3e-2, div_final=10, div=10)\n</code></pre>\n<p>So in English, this is a one cycle learning rate schedule with <strong>30 epochs</strong> starting with lr=1e-4, increasing to 1e-3 and then decreasing back to 1e-4 following cosine anealing schedule. The <strong>loss function</strong> is the good old cross-entropy. Also, a <strong>weight decay</strong> of 3e-2 was used, a <strong>gradient clip</strong> of 0.5 and the train was done with <strong>mixed-precision</strong> so that my GTX 1080 can handle a <strong>batch size</strong> of 32 with 128x1024 image size.</p>\n<p>One training epoch takes about 1 minute on my GTX 1080,  I guess it's not bad considering that I'm doing waveform augmentations on CPU that even with p=0.25 take some time.</p>\n<p><strong>Inference</strong><br>\nThis is the fun part because it was almost by mistake that I realised that making inference with smaller tiles is way better. I presume that this is the case because I'm training with cross-entropy for a single label problem but the test data is labelled with multiple labels. By using smaller crops the predictions are more multilabel friendly. The reason I've been using cross-entropy instead of binary cross-entropy and sigmoid for the typical multilabel problem is that for me the convergence was much faster using the cross-entropy approach and with better results. Maybe I made a mistake somewhere I don't know, I didn't investigate it in much detail.</p>\n<ul>\n<li><p>Run predictions on crops of the spectrogram with a width of 64, 128 and 256 (remember training was done with 1024), calculate the max probability for each class for each case (64, 128, 256) and the average of the 3 cases. The average of the 3 gave me public lb 0.928 on my best single model that I describe above, compared to 0.925 for just the 128 width inference.</p></li>\n<li><p>The final solution with public lb 0.932 and private lb 0.940 is an ensemble of a few training interations with some modifications. (I will update this tomorrow with more information).</p></li>\n</ul>\n<p><strong>dl_pipeline</strong><br>\nAnd again the code for this solution is now public on this repo: <a href=\"https://github.com/mnpinto/dl_pipeline\" target=\"_blank\">https://github.com/mnpinto/dl_pipeline</a></p>\n<p>The following code should correspond to the best single model solution but I need to check if I didn't mess up anything when cleaning the code:</p>\n<pre><code>#!/bin/bash\narch='densenet121'\nmodel_name='model_0'\nsample_rate=32000\nn_mels=128\nhop_length=640\n\nfor fold in 0 1 2 3 4\ndo\n    echo \"Training $model for fold $fold\"\n    kaggle_rainforest2021 --fold $fold --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --bs 32 --head_ps 0.8 \\\n        --tile_width 1024 --mixup true &gt;&gt; log.train\ndone\n\nfor tw in 64 128 256\ndo\n    echo \"Generate predictions for $model with tile_width of $tw\"\n    kaggle_rainforest2021 --run_test true --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --tile_width $tw \\\n        --save_preds true &gt;&gt; log.predict\ndone\n</code></pre>\n<hr>\n<p>Thanks for reading!</p>",
      "rawMarkdown": "**In summary:**\n- Best single model (0.925 public lb): densenet121 features + fastai head\n- Loss function: cross entropy\n- Sampling: 128x1024 crops around true positive samples\n- Spectrogram parameters: n mels 128, hop length 640, sample rate 32000\n- Augmentations: clipping distortion, pitch shift and mixup\n- Inference: Predict on crops with a small width (e.g. 128x128 instead of 128x1024 used for training) and calculate the max probability for each of the 24 classes.\n\n**Introductory monologue**\nFirst and foremost, this was an interesting competition and a good learning opportunity as it is often the case in Kaggle! One “problem” of this competition is that the test data was labeled with a different method and no samples of labeled test data were provided. This makes it difficult to get a sense of validation score and increases the danger of overfiting the public test results. In fact, I almost gave up on this competition when I realized this was the case. But eventually I decided to get back to it and work on a simple solution and on a python library – **dl_pipeline** ( https://github.com/mnpinto/dl_pipeline) – that I will use as a general framework for future kaggle competitions in general. Initially, the idea for dl_pipeline was just to keep my code more organized and more reusable but I figured that maybe there’s also some value in sharing it. \n\n**Data preprocessing**\nSave all wave files in npy files with sample rate of 32000 Hz to save time.\n```python\ndef audio2npy(file, path_save:Path, sample_rate=32_000):\n    path_save.mkdir(exist_ok=True, parents=True)\n    wave, _ = librosa.load(file, sr=sample_rate)\n    np.save(path_save/f'{file.stem}.npy', wave)\n```\nI didn't convert the audio to spectrograms right away since I still want the ability to use audio augmentation on the waveforms.\n\n**Augmentations and Spectrograms**\n- First I create crops on the waveform including the true positive labels with a number of samples calculated so that the spectrogram will have a width of 1024. \n\n**Note:** Cropping before applying the augmentations is much faster than the other way around.\n\n- Then for the waveform augmentations I used the **audiomentations** library (https://github.com/iver56/audiomentations). I ended up using just the following augmentations as based on public lb I didn't find that others were helping, although this would require a proper validation to take any conclusions. \n\n```\ndef audio_augment(sample_rate, p=0.25):\n    return Pipeline([\n        ClippingDistortion(sample_rate, max_percentile_threshold=10, p=p),\n        PitchShift(sample_rate, min_semitones=-8, max_semitones=8, p=p),\n    ])\n```\n\n**Note:** Some augmentations are much slower, for example, pitch shift and time stretch. When using those augmentations the probability of use makes a big difference in how long the training takes. \n\n- Then I searched the fastest way to convert the audio to spectrograms in the GPU and I ended up using **nnAudio** (https://github.com/KinWaiCheuk/nnAudio). Again, converting to spectrogram after the waveform is cropped is a nice gain in processing time.\n\n**Model**\nI tried several models but the one that got me a better result in the public leaderboard was densenet121 and the second-best ResNeSt50. One particularity is that I use for all the models the fastai head with strong dropout.\n\nThe fastai head (using `create_head(num_features*2, num_classes, ps=0.8)`).\n\n```\n(1): Sequential(\n      (0): AdaptiveConcatPool2d(\n        (ap): AdaptiveAvgPool2d(output_size=1)\n        (mp): AdaptiveMaxPool2d(output_size=1)\n      )\n      (1): Flatten(full=False)\n      (2): BatchNorm1d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (3): Dropout(p=0.4, inplace=False)\n      (4): Linear(in_features=2048, out_features=512, bias=False)\n      (5): ReLU(inplace=True)\n      (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (7): Dropout(p=0.8, inplace=False)\n      (8): Linear(in_features=512, out_features=24, bias=False)\n    )\n```\n\n**Training**\nI guess code speaks more than words, particularly for those familiar with fastai:\n\n```\nbs = 32\nlearn = Learner(dls, model, loss_func=cross_entropy, metrics=[accuracy, lrap], cbs=cbs)\nlearn.to_fp16(clip=0.5);\nlearn.fit_one_cycle(30, 1e-3, wd=3e-2, div_final=10, div=10)\n```\n\nSo in English, this is a one cycle learning rate schedule with **30 epochs** starting with lr=1e-4, increasing to 1e-3 and then decreasing back to 1e-4 following cosine anealing schedule. The **loss function** is the good old cross-entropy. Also, a **weight decay** of 3e-2 was used, a **gradient clip** of 0.5 and the train was done with **mixed-precision** so that my GTX 1080 can handle a **batch size** of 32 with 128x1024 image size.\n\nOne training epoch takes about 1 minute on my GTX 1080,  I guess it's not bad considering that I'm doing waveform augmentations on CPU that even with p=0.25 take some time.\n\n**Inference**\nThis is the fun part because it was almost by mistake that I realised that making inference with smaller tiles is way better. I presume that this is the case because I'm training with cross-entropy for a single label problem but the test data is labelled with multiple labels. By using smaller crops the predictions are more multilabel friendly. The reason I've been using cross-entropy instead of binary cross-entropy and sigmoid for the typical multilabel problem is that for me the convergence was much faster using the cross-entropy approach and with better results. Maybe I made a mistake somewhere I don't know, I didn't investigate it in much detail.\n\n- Run predictions on crops of the spectrogram with a width of 64, 128 and 256 (remember training was done with 1024), calculate the max probability for each class for each case (64, 128, 256) and the average of the 3 cases. The average of the 3 gave me public lb 0.928 on my best single model that I describe above, compared to 0.925 for just the 128 width inference.\n\n- The final solution with public lb 0.932 and private lb 0.940 is an ensemble of a few training interations with some modifications. (I will update this tomorrow with more information).\n\n**dl_pipeline**\nAnd again the code for this solution is now public on this repo: https://github.com/mnpinto/dl_pipeline\n\nThe following code should correspond to the best single model solution but I need to check if I didn't mess up anything when cleaning the code:\n\n```\n#!/bin/bash\narch='densenet121'\nmodel_name='model_0'\nsample_rate=32000\nn_mels=128\nhop_length=640\n\nfor fold in 0 1 2 3 4\ndo\n    echo \"Training $model for fold $fold\"\n    kaggle_rainforest2021 --fold $fold --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --bs 32 --head_ps 0.8 \\\n        --tile_width 1024 --mixup true >> log.train\ndone\n\nfor tw in 64 128 256\ndo\n    echo \"Generate predictions for $model with tile_width of $tw\"\n    kaggle_rainforest2021 --run_test true --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --tile_width $tw \\\n        --save_preds true >> log.predict\ndone\n```\n----------------------------\nThanks for reading!",
      "votes": null
    },
    {
      "id": "1207680",
      "postDate": "02/18/2021 01:02:06",
      "content": "<p>Congrats on 29th solo place and thanks for sharing solution <a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a> </p>",
      "rawMarkdown": "Congrats on 29th solo place and thanks for sharing solution @mnpinto",
      "votes": null
    },
    {
      "id": "1208503",
      "postDate": "02/18/2021 10:00:22",
      "content": "<p><a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a> Congrats on your achievement. Yes it seems simple and straight forward. Your code speaks good. Thanks for your writeup.</p>",
      "rawMarkdown": "mnpinto Congrats on your achievement. Yes it seems simple and straight forward. Your code speaks good. Thanks for your writeup.",
      "votes": null
    },
    {
      "id": "1209274",
      "postDate": "02/18/2021 19:41:49",
      "content": "<p>Congratz Miguel, and thanks for open-sourcing your code :)</p>",
      "rawMarkdown": "Congratz Miguel, and thanks for open-sourcing your code :)",
      "votes": null
    },
    {
      "id": "1209461",
      "postDate": "02/18/2021 23:03:12",
      "content": "<p>Thanks for sharing.  I'll ask you this question but it is valid for all those who did not use the FP data.  TP only gives you ground truth of 1.  What ground truth 0 did you all use?</p>\n<p>I am quite puzzled to be honest.</p>\n<p>Did you assume that the ground truth is zero for all other species in each TP crop? If so then it is actually wrong,  Some crops have more than one species.</p>\n<p>Anyway, I'd be happy to hear what you did for ground truth.  And congrats on the good result.</p>",
      "rawMarkdown": "Thanks for sharing.  I'll ask you this question but it is valid for all those who did not use the FP data.  TP only gives you ground truth of 1.  What ground truth 0 did you all use?\n\nI am quite puzzled to be honest.\n\nDid you assume that the ground truth is zero for all other species in each TP crop? If so then it is actually wrong,  Some crops have more than one species.\n\nAnyway, I'd be happy to hear what you did for ground truth.  And congrats on the good result.",
      "votes": null
    },
    {
      "id": "1209499",
      "postDate": "02/18/2021 23:47:17",
      "content": "<p>You are correct <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, I considered 0 to all other species. I worked on this as a single class classification problem using cross-entropy loss for training. Initially I was using smaller audio crops (128x128 or 128x256) and the rationale was that probably there's not much overlap of classes in the small crops around the TPs (i.e. if I know class \"A\" is observed in that small crop, it's quite likely that most, if any, of the other 23 won't be in the same crop). And indeed with cross-entropy loss the model converges quite well. I now see that the idea of using BCE with TP as 1 and FP as 0 and masking all other values is what I was missing and a great way to incorporate the FP. Nevertheless, with Chris Deotte post-processing my best single model gets 0.950 on private LB (resnest50 and resnest101), generating the predictions over crops of 128 (steps 64) and 256 (steps 128). </p>",
      "rawMarkdown": "You are correct @cpmpml, I considered 0 to all other species. I worked on this as a single class classification problem using cross-entropy loss for training. Initially I was using smaller audio crops (128x128 or 128x256) and the rationale was that probably there's not much overlap of classes in the small crops around the TPs (i.e. if I know class \"A\" is observed in that small crop, it's quite likely that most, if any, of the other 23 won't be in the same crop). And indeed with cross-entropy loss the model converges quite well. I now see that the idea of using BCE with TP as 1 and FP as 0 and masking all other values is what I was missing and a great way to incorporate the FP. Nevertheless, with Chris Deotte post-processing my best single model gets 0.950 on private LB (resnest50 and resnest101), generating the predictions over crops of 128 (steps 64) and 256 (steps 128).",
      "votes": null
    },
    {
      "id": "1329963",
      "postDate": "05/31/2021 14:01:50",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a> , may I ask what's the speed of nnAudio compared to torchlibrosa? Looks they both use conv1d to do STFT, and how's the performance of nnAudio, any degrade compare to librosa or torchlibrosa?</p>",
      "rawMarkdown": "Thanks for sharing @mnpinto , may I ask what's the speed of nnAudio compared to torchlibrosa? Looks they both use conv1d to do STFT, and how's the performance of nnAudio, any degrade compare to librosa or torchlibrosa?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207680,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 01:02:06",
      "content": "<p>Congrats on 29th solo place and thanks for sharing solution <a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208503,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 10:00:22",
      "content": "<p><a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a> Congrats on your achievement. Yes it seems simple and straight forward. Your code speaks good. Thanks for your writeup.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209274,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 19:41:49",
      "content": "<p>Congratz Miguel, and thanks for open-sourcing your code :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209461,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 23:03:12",
      "content": "<p>Thanks for sharing.  I'll ask you this question but it is valid for all those who did not use the FP data.  TP only gives you ground truth of 1.  What ground truth 0 did you all use?</p>\n<p>I am quite puzzled to be honest.</p>\n<p>Did you assume that the ground truth is zero for all other species in each TP crop? If so then it is actually wrong,  Some crops have more than one species.</p>\n<p>Anyway, I'd be happy to hear what you did for ground truth.  And congrats on the good result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209499,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "02/18/2021 23:47:17",
          "content": "<p>You are correct <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, I considered 0 to all other species. I worked on this as a single class classification problem using cross-entropy loss for training. Initially I was using smaller audio crops (128x128 or 128x256) and the rationale was that probably there's not much overlap of classes in the small crops around the TPs (i.e. if I know class \"A\" is observed in that small crop, it's quite likely that most, if any, of the other 23 won't be in the same crop). And indeed with cross-entropy loss the model converges quite well. I now see that the idea of using BCE with TP as 1 and FP as 0 and masking all other values is what I was missing and a great way to incorporate the FP. Nevertheless, with Chris Deotte post-processing my best single model gets 0.950 on private LB (resnest50 and resnest101), generating the predictions over crops of 128 (steps 64) and 256 (steps 128). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1329963,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "05/31/2021 14:01:50",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a> , may I ask what's the speed of nnAudio compared to torchlibrosa? Looks they both use conv1d to do STFT, and how's the performance of nnAudio, any degrade compare to librosa or torchlibrosa?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207597": "**In summary:**\n- Best single model (0.925 public lb): densenet121 features + fastai head\n- Loss function: cross entropy\n- Sampling: 128x1024 crops around true positive samples\n- Spectrogram parameters: n mels 128, hop length 640, sample rate 32000\n- Augmentations: clipping distortion, pitch shift and mixup\n- Inference: Predict on crops with a small width (e.g. 128x128 instead of 128x1024 used for training) and calculate the max probability for each of the 24 classes.\n\n**Introductory monologue**\nFirst and foremost, this was an interesting competition and a good learning opportunity as it is often the case in Kaggle! One “problem” of this competition is that the test data was labeled with a different method and no samples of labeled test data were provided. This makes it difficult to get a sense of validation score and increases the danger of overfiting the public test results. In fact, I almost gave up on this competition when I realized this was the case. But eventually I decided to get back to it and work on a simple solution and on a python library – **dl_pipeline** ( https://github.com/mnpinto/dl_pipeline) – that I will use as a general framework for future kaggle competitions in general. Initially, the idea for dl_pipeline was just to keep my code more organized and more reusable but I figured that maybe there’s also some value in sharing it. \n\n**Data preprocessing**\nSave all wave files in npy files with sample rate of 32000 Hz to save time.\n```python\ndef audio2npy(file, path_save:Path, sample_rate=32_000):\n    path_save.mkdir(exist_ok=True, parents=True)\n    wave, _ = librosa.load(file, sr=sample_rate)\n    np.save(path_save/f'{file.stem}.npy', wave)\n```\nI didn't convert the audio to spectrograms right away since I still want the ability to use audio augmentation on the waveforms.\n\n**Augmentations and Spectrograms**\n- First I create crops on the waveform including the true positive labels with a number of samples calculated so that the spectrogram will have a width of 1024. \n\n**Note:** Cropping before applying the augmentations is much faster than the other way around.\n\n- Then for the waveform augmentations I used the **audiomentations** library (https://github.com/iver56/audiomentations). I ended up using just the following augmentations as based on public lb I didn't find that others were helping, although this would require a proper validation to take any conclusions. \n\n```\ndef audio_augment(sample_rate, p=0.25):\n    return Pipeline([\n        ClippingDistortion(sample_rate, max_percentile_threshold=10, p=p),\n        PitchShift(sample_rate, min_semitones=-8, max_semitones=8, p=p),\n    ])\n```\n\n**Note:** Some augmentations are much slower, for example, pitch shift and time stretch. When using those augmentations the probability of use makes a big difference in how long the training takes. \n\n- Then I searched the fastest way to convert the audio to spectrograms in the GPU and I ended up using **nnAudio** (https://github.com/KinWaiCheuk/nnAudio). Again, converting to spectrogram after the waveform is cropped is a nice gain in processing time.\n\n**Model**\nI tried several models but the one that got me a better result in the public leaderboard was densenet121 and the second-best ResNeSt50. One particularity is that I use for all the models the fastai head with strong dropout.\n\nThe fastai head (using `create_head(num_features*2, num_classes, ps=0.8)`).\n\n```\n(1): Sequential(\n      (0): AdaptiveConcatPool2d(\n        (ap): AdaptiveAvgPool2d(output_size=1)\n        (mp): AdaptiveMaxPool2d(output_size=1)\n      )\n      (1): Flatten(full=False)\n      (2): BatchNorm1d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (3): Dropout(p=0.4, inplace=False)\n      (4): Linear(in_features=2048, out_features=512, bias=False)\n      (5): ReLU(inplace=True)\n      (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (7): Dropout(p=0.8, inplace=False)\n      (8): Linear(in_features=512, out_features=24, bias=False)\n    )\n```\n\n**Training**\nI guess code speaks more than words, particularly for those familiar with fastai:\n\n```\nbs = 32\nlearn = Learner(dls, model, loss_func=cross_entropy, metrics=[accuracy, lrap], cbs=cbs)\nlearn.to_fp16(clip=0.5);\nlearn.fit_one_cycle(30, 1e-3, wd=3e-2, div_final=10, div=10)\n```\n\nSo in English, this is a one cycle learning rate schedule with **30 epochs** starting with lr=1e-4, increasing to 1e-3 and then decreasing back to 1e-4 following cosine anealing schedule. The **loss function** is the good old cross-entropy. Also, a **weight decay** of 3e-2 was used, a **gradient clip** of 0.5 and the train was done with **mixed-precision** so that my GTX 1080 can handle a **batch size** of 32 with 128x1024 image size.\n\nOne training epoch takes about 1 minute on my GTX 1080,  I guess it's not bad considering that I'm doing waveform augmentations on CPU that even with p=0.25 take some time.\n\n**Inference**\nThis is the fun part because it was almost by mistake that I realised that making inference with smaller tiles is way better. I presume that this is the case because I'm training with cross-entropy for a single label problem but the test data is labelled with multiple labels. By using smaller crops the predictions are more multilabel friendly. The reason I've been using cross-entropy instead of binary cross-entropy and sigmoid for the typical multilabel problem is that for me the convergence was much faster using the cross-entropy approach and with better results. Maybe I made a mistake somewhere I don't know, I didn't investigate it in much detail.\n\n- Run predictions on crops of the spectrogram with a width of 64, 128 and 256 (remember training was done with 1024), calculate the max probability for each class for each case (64, 128, 256) and the average of the 3 cases. The average of the 3 gave me public lb 0.928 on my best single model that I describe above, compared to 0.925 for just the 128 width inference.\n\n- The final solution with public lb 0.932 and private lb 0.940 is an ensemble of a few training interations with some modifications. (I will update this tomorrow with more information).\n\n**dl_pipeline**\nAnd again the code for this solution is now public on this repo: https://github.com/mnpinto/dl_pipeline\n\nThe following code should correspond to the best single model solution but I need to check if I didn't mess up anything when cleaning the code:\n\n```\n#!/bin/bash\narch='densenet121'\nmodel_name='model_0'\nsample_rate=32000\nn_mels=128\nhop_length=640\n\nfor fold in 0 1 2 3 4\ndo\n    echo \"Training $model for fold $fold\"\n    kaggle_rainforest2021 --fold $fold --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --bs 32 --head_ps 0.8 \\\n        --tile_width 1024 --mixup true >> log.train\ndone\n\nfor tw in 64 128 256\ndo\n    echo \"Generate predictions for $model with tile_width of $tw\"\n    kaggle_rainforest2021 --run_test true --model_name $model_name \\\n        --model $arch --sample_rate $sample_rate --n_mels $n_mels \\\n        --hop_length $hop_length --tile_width $tw \\\n        --save_preds true >> log.predict\ndone\n```\n----------------------------\nThanks for reading!",
    "1207680": "Congrats on 29th solo place and thanks for sharing solution @mnpinto",
    "1208503": "mnpinto Congrats on your achievement. Yes it seems simple and straight forward. Your code speaks good. Thanks for your writeup.",
    "1209274": "Congratz Miguel, and thanks for open-sourcing your code :)",
    "1209461": "Thanks for sharing.  I'll ask you this question but it is valid for all those who did not use the FP data.  TP only gives you ground truth of 1.  What ground truth 0 did you all use?\n\nI am quite puzzled to be honest.\n\nDid you assume that the ground truth is zero for all other species in each TP crop? If so then it is actually wrong,  Some crops have more than one species.\n\nAnyway, I'd be happy to hear what you did for ground truth.  And congrats on the good result.",
    "1209499": "You are correct @cpmpml, I considered 0 to all other species. I worked on this as a single class classification problem using cross-entropy loss for training. Initially I was using smaller audio crops (128x128 or 128x256) and the rationale was that probably there's not much overlap of classes in the small crops around the TPs (i.e. if I know class \"A\" is observed in that small crop, it's quite likely that most, if any, of the other 23 won't be in the same crop). And indeed with cross-entropy loss the model converges quite well. I now see that the idea of using BCE with TP as 1 and FP as 0 and masking all other values is what I was missing and a great way to incorporate the FP. Nevertheless, with Chris Deotte post-processing my best single model gets 0.950 on private LB (resnest50 and resnest101), generating the predictions over crops of 128 (steps 64) and 256 (steps 128).",
    "1329963": "Thanks for sharing @mnpinto , may I ask what's the speed of nnAudio compared to torchlibrosa? Looks they both use conv1d to do STFT, and how's the performance of nnAudio, any degrade compare to librosa or torchlibrosa?"
  },
  "source": "meta"
}