{
  "id": 234154,
  "title": "Fast training and precomputed melspecs images",
  "url": "/competitions/birdclef-2021/discussion/234154",
  "author_name": "kkiller",
  "post_date": "2021-04-22T23:14:50.894000",
  "votes": 63,
  "comment_count": 57,
  "views": 0,
  "content": "<p>My training pipeline was used to last 9 hours for just 12 epochs. This duration has been incredibly reduced when I precomptued the mels and converted them into \"<strong>np.uint8</strong>\" dtypes (one of the smallest  numpy dtypes). Now I can train a <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab\" target=\"_blank\"> full model on Kaggle</a> and one fold of 20 epochs could last <strong>less than two hours</strong>.</p>\n<p>Here are the links to those handy datasets:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-2021\" target=\"_blank\">7 seconds records numpy images with truncation</a></li>\n</ul>\n<p>If you're interrested in the whole record melspecs images (no truncation):</p>\n<p><a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part1\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part1</a><br>\n<a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part2\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part2</a><br>\n<a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part3\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part3</a><br>\n<a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part4\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part4</a></p>",
  "messages": [
    {
      "id": 1281413,
      "postDate": "2021-04-22T23:14:50.893Z",
      "content": "<p>My training pipeline was used to last 9 hours for just 12 epochs. This duration has been incredibly reduced when I precomptued the mels and converted them into \"<strong>np.uint8</strong>\" dtypes (one of the smallest  numpy dtypes). Now I can train a <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab\" target=\"_blank\"> full model on Kaggle</a> and one fold of 20 epochs could last <strong>less than two hours</strong>.</p>\n<p>Here are the links to those handy datasets:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-2021\" target=\"_blank\">7 seconds records numpy images with truncation</a></li>\n</ul>\n<p>If you're interrested in the whole record melspecs images (no truncation):</p>\n<p><a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part1\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part1</a><br>\n<a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part2\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part2</a><br>\n<a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part3\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part3</a><br>\n<a href=\"https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part4\" target=\"_blank\">https://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part4</a></p>",
      "rawMarkdown": "My training pipeline was used to last 9 hours for just 12 epochs. This duration has been incredibly reduced when I precomptued the mels and converted them into \"**np.uint8**\" dtypes (one of the smallest  numpy dtypes). Now I can train a [ full model on Kaggle](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab) and one fold of 20 epochs could last **less than two hours**.\n\nHere are the links to those handy datasets:\n\n* [7 seconds records numpy images with truncation](https://www.kaggle.com/kneroma/kkiller-birdclef-2021)\n\nIf you're interrested in the whole record melspecs images (no truncation):\n\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part1\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part2\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part3\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part4\n",
      "votes": 60
    },
    {
      "id": 1285218,
      "postDate": "2021-04-26T17:40:29.590Z",
      "content": "<p>In case anyone's using tensorflow datasets:</p>\n<p>I recently (re-)ran some performance experiments with a desktop GPU, and found it fastest to apply time-domain augmentations in the TF dataset (which runs with heavily parallelization on the machine's CPU), and then apply the STFT+MelSpec extraction in the model code (which runs on the GPU: The STFT+Melspec ops are basically just batched convolution and matmul, which is exactly whet the GPU is best at). </p>\n<p>Producing a batch of 32 time-domain augmented audio segments now takes ~0.08 seconds. When the melspec operations are included in the dataset it takes about 0.4 seconds to produce a batch.  The time domain augmentations include mix-up addition of two examples, gain randomization, and noise addition. I've also got a few melspec-domain augmentations, including emulating a random low-pass filter in the mel domain, which happen in the model graph now.</p>\n<p>Meanwhile, once the melspec computation is in the model graph, the model trains at ~0.23s per batch, so isn't input bound. Obvs, all the perf numbers will depend on your particular machine and model architecture; the point is that if you can keep the model well-fed with time-domain data, it may be faster to make the melspec part of the model so it runs on GPU/TPU.</p>",
      "rawMarkdown": "In case anyone's using tensorflow datasets:\n\nI recently (re-)ran some performance experiments with a desktop GPU, and found it fastest to apply time-domain augmentations in the TF dataset (which runs with heavily parallelization on the machine's CPU), and then apply the STFT+MelSpec extraction in the model code (which runs on the GPU: The STFT+Melspec ops are basically just batched convolution and matmul, which is exactly whet the GPU is best at). \n\nProducing a batch of 32 time-domain augmented audio segments now takes ~0.08 seconds. When the melspec operations are included in the dataset it takes about 0.4 seconds to produce a batch.  The time domain augmentations include mix-up addition of two examples, gain randomization, and noise addition. I've also got a few melspec-domain augmentations, including emulating a random low-pass filter in the mel domain, which happen in the model graph now.\n\nMeanwhile, once the melspec computation is in the model graph, the model trains at ~0.23s per batch, so isn't input bound. Obvs, all the perf numbers will depend on your particular machine and model architecture; the point is that if you can keep the model well-fed with time-domain data, it may be faster to make the melspec part of the model so it runs on GPU/TPU.",
      "votes": 9,
      "replies": [
        {
          "id": 1290385,
          "postDate": "2021-05-01T22:18:11.227Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> for your suggestions. For pytorch users who want to compute the mels on GPU, there are <a href=\"https://pytorch.org/audio/stable/index.html\" target=\"_blank\">torchaudio</a> and <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">torch-audiomentations</a> .</p>",
          "rawMarkdown": "Thanks @tomdenton for your suggestions. For pytorch users who want to compute the mels on GPU, there are [torchaudio](https://pytorch.org/audio/stable/index.html) and [torch-audiomentations](https://github.com/asteroid-team/torch-audiomentations) .",
          "votes": 7
        }
      ]
    },
    {
      "id": 1282331,
      "postDate": "2021-04-23T19:44:45.030Z",
      "content": "<p>Thanks for sharing this technique! I'm wondering if converting mels to np.uint8 would result in the loss of information, which would then negatively impact the model. Is this a trade off between training time and model accuracy?</p>",
      "rawMarkdown": "Thanks for sharing this technique! I'm wondering if converting mels to np.uint8 would result in the loss of information, which would then negatively impact the model. Is this a trade off between training time and model accuracy?",
      "votes": 5,
      "replies": [
        {
          "id": 1282466,
          "postDate": "2021-04-24T01:11:17.337Z",
          "content": "<p>A good point over there. But, for me, even without the pre-computation, I was using cv2 images so it does not change my pipeline. I won't expect the information loss to be too much damaging.</p>",
          "rawMarkdown": "A good point over there. But, for me, even without the pre-computation, I was using cv2 images so it does not change my pipeline. I won't expect the information loss to be too much damaging.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1311108,
      "postDate": "2021-05-17T06:53:33.883Z",
      "content": "<p>Hi, your notebook is great for novice like me. I have a small question, in <a href=\"https://www.kaggle.com/kneroma/birdclef-mels-computer-public\" target=\"_blank\">BirdCLEF Mels Computer [Public]\n</a> you call <code>image = mono_to_color(melspec)</code> without mean and std. </p>\n<pre><code>def mono_to_color(X, eps=1e-6, mean=None, std=None):\n    mean = mean or X.mean()\n    std = std or X.std()\n    X = (X - mean) / (std + eps)\n\n    _min, _max = X.min(), X.max()\n\n    if (_max - _min) &gt; eps:\n        V = np.clip(X, _min, _max)\n        V = 255 * (V - _min) / (_max - _min)\n        V = V.astype(np.uint8)\n    else:\n        V = np.zeros_like(X, dtype=np.uint8)\n\n    return V\n</code></pre>\n<p>This is the first time for me to handle audio data, I find some <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/177551\" target=\"_blank\">discussion</a> that normalize in another way. It it ok to normalize with self's mean and std?</p>",
      "rawMarkdown": "Hi, your notebook is great for novice like me. I have a small question, in [BirdCLEF Mels Computer [Public]\n](https://www.kaggle.com/kneroma/birdclef-mels-computer-public) you call `image = mono_to_color(melspec)` without mean and std. \n```\ndef mono_to_color(X, eps=1e-6, mean=None, std=None):\n    mean = mean or X.mean()\n    std = std or X.std()\n    X = (X - mean) / (std + eps)\n    \n    _min, _max = X.min(), X.max()\n\n    if (_max - _min) > eps:\n        V = np.clip(X, _min, _max)\n        V = 255 * (V - _min) / (_max - _min)\n        V = V.astype(np.uint8)\n    else:\n        V = np.zeros_like(X, dtype=np.uint8)\n\n    return V\n```\nThis is the first time for me to handle audio data, I find some [discussion](https://www.kaggle.com/c/birdsong-recognition/discussion/177551) that normalize in another way. It it ok to normalize with self's mean and std?",
      "votes": 4,
      "replies": [
        {
          "id": 1312257,
          "postDate": "2021-05-18T00:15:22.847Z",
          "content": "<p>Great question.</p>\n<p>I've also seen theo's discussion. Personally I'm not fan of carrying ImageNet's mean and std everywhere for normalization purposes. Here, we're not doing a simple transfer learning (the model will be trained on a completely different dataset for many and many epochs) and the impact of the pretrained weights is questionable. </p>\n<p>I would say It's much a question of taste. I would be happy to read someone's results after moving from the original <strong>mono_to_color</strong> to the new one.</p>\n<p>The fact that the means and stds are not constant over the samples is actually a feature and not a bug as the mels are not true images. A mel is semantically the same when divided by a constant or not, same for addition. Some people are doing even more complicated self transformations, like raising the whole mels into the power of 3.</p>",
          "rawMarkdown": "Great question.\n\nI've also seen theo's discussion. Personally I'm not fan of carrying ImageNet's mean and std everywhere for normalization purposes. Here, we're not doing a simple transfer learning (the model will be trained on a completely different dataset for many and many epochs) and the impact of the pretrained weights is questionable. \n\nI would say It's much a question of taste. I would be happy to read someone's results after moving from the original **mono_to_color** to the new one.\n\nThe fact that the means and stds are not constant over the samples is actually a feature and not a bug as the mels are not true images. A mel is semantically the same when divided by a constant or not, same for addition. Some people are doing even more complicated self transformations, like raising the whole mels into the power of 3.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1310767,
      "postDate": "2021-05-16T23:51:52.803Z",
      "content": "<p>I \"ran all\" and got an error at the last step. Seemed like resnest wasn't working, because I then tried a resnet50 and that is working fine. Any ideas why the last code bubble would throw an error?</p>\n<p>Thanks in advance for the help !</p>\n<p>The error specifically was that the Resnest50 downloader only got to 5/6 and then failed, yet the exception error was not thrown for some reason</p>",
      "rawMarkdown": "I \"ran all\" and got an error at the last step. Seemed like resnest wasn't working, because I then tried a resnet50 and that is working fine. Any ideas why the last code bubble would throw an error?\n\nThanks in advance for the help !\n\nThe error specifically was that the Resnest50 downloader only got to 5/6 and then failed, yet the exception error was not thrown for some reason\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1312266,
          "postDate": "2021-05-18T00:27:29.853Z",
          "content": "<p>I think the resnest50 pretrained weights' link is broken.</p>\n<p>Please refer to <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/comments#1303631\" target=\"_blank\">this thread</a> for a workaround.</p>",
          "rawMarkdown": "I think the resnest50 pretrained weights' link is broken.\n\nPlease refer to [this thread](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/comments#1303631) for a workaround.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1305644,
      "postDate": "2021-05-13T12:06:18.607Z",
      "content": "<p>Thanks for your sharing! I also have a question, are data augmentations like Pink noise/Gaussian noise/Gaussian SNR/… applied before mels being computed?</p>",
      "rawMarkdown": "Thanks for your sharing! I also have a question, are data augmentations like Pink noise/Gaussian noise/Gaussian SNR/... applied before mels being computed?",
      "votes": 1,
      "replies": [
        {
          "id": 1313111,
          "postDate": "2021-05-18T12:10:31.937Z",
          "content": "<p>No. Still you can use spec augs. If you're interested in raw audio augs, precomputed mels are not a good choice.</p>",
          "rawMarkdown": "No. Still you can use spec augs. If you're interested in raw audio augs, precomputed mels are not a good choice.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1290016,
      "postDate": "2021-05-01T14:56:23.903Z",
      "content": "<p>Thank you for sharing all this, I am learning a lot from you!</p>",
      "rawMarkdown": "Thank you for sharing all this, I am learning a lot from you!",
      "votes": 1,
      "replies": [
        {
          "id": 1290386,
          "postDate": "2021-05-01T22:18:56.080Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/daniellga\" target=\"_blank\">@daniellga</a> </p>",
          "rawMarkdown": "Thanks @daniellga ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1289816,
      "postDate": "2021-05-01T11:54:00.690Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>, thanks for sharing these valuable resources! I had a few questions. I've joined this competition recently and it's my first time working with audio, so sorry in advance if my questions are stupid :)</p>\n<p>1- Doesn't it harm the model when it gets trained on random 7 seconds of the recording where we do not know for sure that if there is any bird present? I mean the randomly selected 7 sec can be from a part of the recording with pure noise and the model sees a real bird label for that noise. Of course some of these 7 seconds do have the bird but what about those that do not? And if I'm correct in the way I see the problem, are there alternatives to your strategy?</p>\n<p>2- In the notebook that you shared which builds these datasets, in the last function you are setting the <code>step</code> variable to DURATION * 0.666 * SR. What is that 0.666 about? Does it mean that the final separated records will be less than 7 secs?</p>\n<p>3- For this last one I do not really expect an answer and answer this if you're okay with sharing: Is your LB (or near that) has been achieved using these datasets?</p>",
      "rawMarkdown": "Hey @kneroma, thanks for sharing these valuable resources! I had a few questions. I've joined this competition recently and it's my first time working with audio, so sorry in advance if my questions are stupid :)\n\n1- Doesn't it harm the model when it gets trained on random 7 seconds of the recording where we do not know for sure that if there is any bird present? I mean the randomly selected 7 sec can be from a part of the recording with pure noise and the model sees a real bird label for that noise. Of course some of these 7 seconds do have the bird but what about those that do not? And if I'm correct in the way I see the problem, are there alternatives to your strategy?\n\n2- In the notebook that you shared which builds these datasets, in the last function you are setting the `step` variable to DURATION * 0.666 * SR. What is that 0.666 about? Does it mean that the final separated records will be less than 7 secs?\n\n3- For this last one I do not really expect an answer and answer this if you're okay with sharing: Is your LB (or near that) has been achieved using these datasets?",
      "votes": 1,
      "replies": [
        {
          "id": 1290379,
          "postDate": "2021-05-01T22:15:05.683Z",
          "content": "<ol>\n<li><p>The main problem is that the train set is weakly labeled (target is only available for the whole clip, not for audio segments). Random  chunk is a very basic startegy, and I won't expect it to be the smartest one. There may be several alternatives. For instance, you may train a <strong>call / no call</strong> classifier and afterward consider only the chunks with a call …</p></li>\n<li><p>0.666 is meaningless, I choose this value just to avoid a memory overflow on Kaggle. Generally, if <strong>step &lt; DURATION*SR</strong>, it means that there is an overlapping between your chunks. Depending on how much memory do you have, you can lower this 0.666 to get more overlapped chunks.</p></li>\n<li><p>My first subs was exclusively based on those datasets (I got 0.73 LB with them). Currently, I'm experimenting many ideas locally, with no great succes xD .</p></li>\n</ol>",
          "rawMarkdown": "1. The main problem is that the train set is weakly labeled (target is only available for the whole clip, not for audio segments). Random  chunk is a very basic startegy, and I won't expect it to be the smartest one. There may be several alternatives. For instance, you may train a **call / no call** classifier and afterward consider only the chunks with a call ...\n\n2. 0.666 is meaningless, I choose this value just to avoid a memory overflow on Kaggle. Generally, if **step < DURATION*SR**, it means that there is an overlapping between your chunks. Depending on how much memory do you have, you can lower this 0.666 to get more overlapped chunks.\n\n3. My first subs was exclusively based on those datasets (I got 0.73 LB with them). Currently, I'm experimenting many ideas locally, with no great succes xD .",
          "votes": 7
        },
        {
          "id": 1290525,
          "postDate": "2021-05-02T05:00:59.540Z",
          "content": "<p>Thank you very much!</p>",
          "rawMarkdown": "Thank you very much!",
          "votes": 1
        },
        {
          "id": 1291361,
          "postDate": "2021-05-03T01:15:10.237Z",
          "content": "<p>I tried using an onset analyzer to break it apart but it turned out breaking up the chickadee call into small little segments. <br>\nCode examples here:<br>\n<a href=\"https://github.com/thesteve0/birdclef21/commit/8f7e0e129243bfdcc1f675c007ccd837b60aaa24?branch=8f7e0e129243bfdcc1f675c007ccd837b60aaa24&amp;diff=unified\" target=\"_blank\">https://github.com/thesteve0/birdclef21/commit/8f7e0e129243bfdcc1f675c007ccd837b60aaa24?branch=8f7e0e129243bfdcc1f675c007ccd837b60aaa24&amp;diff=unified</a></p>",
          "rawMarkdown": "I tried using an onset analyzer to break it apart but it turned out breaking up the chickadee call into small little segments. \nCode examples here:\nhttps://github.com/thesteve0/birdclef21/commit/8f7e0e129243bfdcc1f675c007ccd837b60aaa24?branch=8f7e0e129243bfdcc1f675c007ccd837b60aaa24&diff=unified\n"
        },
        {
          "id": 1307648,
          "postDate": "2021-05-14T15:24:41.260Z",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  how much was ur local cv that gave lb 0.73 , is that five fold score or single fold thanks in advance</p>",
          "rawMarkdown": "@kneroma  how much was ur local cv that gave lb 0.73 , is that five fold score or single fold thanks in advance"
        }
      ]
    },
    {
      "id": 1283391,
      "postDate": "2021-04-24T22:07:57.213Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 1283465,
          "postDate": "2021-04-25T00:29:01.730Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> </p>",
          "rawMarkdown": "Thanks @atamazian ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1281847,
      "postDate": "2021-04-23T11:10:12.190Z",
      "content": "<p>Thank for sharing <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> - that's some kinda of <code>kkiller-guns</code> :) very generous of you! <br>\nI wish you'll get what you deserved at the end of the competition</p>",
      "rawMarkdown": "Thank for sharing @kneroma - that's some kinda of `kkiller-guns` :) very generous of you! \nI wish you'll get what you deserved at the end of the competition",
      "votes": 1,
      "replies": [
        {
          "id": 1282426,
          "postDate": "2021-04-23T23:15:40.773Z",
          "content": "<p>Thanks for the kind words <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> </p>",
          "rawMarkdown": "Thanks for the kind words @imeintanis ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1281431,
      "postDate": "2021-04-22T23:50:57.827Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>! That's very useful =)</p>\n<p>It looks like the \"7 seconds records numpy images with truncation\" link is broken, though</p>",
      "rawMarkdown": "Thanks @kneroma! That's very useful =)\n\nIt looks like the \"7 seconds records numpy images with truncation\" link is broken, though",
      "votes": 1,
      "replies": [
        {
          "id": 1281438,
          "postDate": "2021-04-23T00:03:55.503Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/adriel\" target=\"_blank\">@adriel</a> . Link fixed !</p>",
          "rawMarkdown": "Thanks @adriel . Link fixed !",
          "votes": 1
        }
      ]
    },
    {
      "id": 1304809,
      "postDate": "2021-05-12T22:36:28.690Z",
      "content": "<p>This is great, I have a question, is there any advantage if we convert the mel spectrogram directly to a tensor?</p>",
      "rawMarkdown": "This is great, I have a question, is there any advantage if we convert the mel spectrogram directly to a tensor?",
      "votes": 2,
      "replies": [
        {
          "id": 1304988,
          "postDate": "2021-05-13T03:49:07.587Z",
          "content": "<p>The only advantage I can see is that it could avoid the numpy - torch conversion. But since pytorch tensor and numpy arrays can share the same storage, the overhead brings by the conversion step may be negligible.</p>",
          "rawMarkdown": "The only advantage I can see is that it could avoid the numpy - torch conversion. But since pytorch tensor and numpy arrays can share the same storage, the overhead brings by the conversion step may be negligible.",
          "votes": 1
        },
        {
          "id": 1306230,
          "postDate": "2021-05-13T17:15:53.843Z",
          "content": "<p>Quick question, is it possible to use the preprocessed dataset in npy on the submission? </p>",
          "rawMarkdown": "Quick question, is it possible to use the preprocessed dataset in npy on the submission? "
        },
        {
          "id": 1313115,
          "postDate": "2021-05-18T12:12:26.080Z",
          "content": "<p>We don't have access to the hidden test set, except inside the \"submission\" environment.</p>",
          "rawMarkdown": "We don't have access to the hidden test set, except inside the \"submission\" environment."
        }
      ]
    },
    {
      "id": 1304286,
      "postDate": "2021-05-12T14:34:12.573Z",
      "content": "<p>I am learning a lot from you too! Thanks for your sharing.It's very useful for me.</p>",
      "rawMarkdown": "I am learning a lot from you too! Thanks for your sharing.It's very useful for me.",
      "votes": 2,
      "replies": [
        {
          "id": 1304319,
          "postDate": "2021-05-12T14:52:34.347Z",
          "content": "<p>Thanks  <a href=\"https://www.kaggle.com/majunfu\" target=\"_blank\">@majunfu</a> </p>",
          "rawMarkdown": "Thanks  @majunfu ",
          "votes": 1
        },
        {
          "id": 1305695,
          "postDate": "2021-05-13T12:34:46.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  thanku for resources.<br>\nOne ask. <br>\nDuring inference i see there 20 test simulating ogg files broken into 120 samples each lasting 5 sec.<br>\nDoes a  each datafile broken into 120 samples is either no call or call  from a single bird </p>",
          "rawMarkdown": " @kneroma  thanku for resources.\nOne ask. \nDuring inference i see there 20 test simulating ogg files broken into 120 samples each lasting 5 sec.\nDoes a  each datafile broken into 120 samples is either no call or call  from a single bird "
        },
        {
          "id": 1307580,
          "postDate": "2021-05-14T14:34:23.073Z",
          "content": "<p>Hi,I want to know if your model is also obtained through current image training?</p>",
          "rawMarkdown": "Hi,I want to know if your model is also obtained through current image training?"
        }
      ]
    },
    {
      "id": 1318909,
      "postDate": "2021-05-22T17:03:24.760Z",
      "content": "<p>thanks for sharing, one question, what's the difference between truncation and no truncation?</p>",
      "rawMarkdown": "thanks for sharing, one question, what's the difference between truncation and no truncation?"
    },
    {
      "id": 1313614,
      "postDate": "2021-05-18T16:14:54.223Z",
      "content": "<p>Curious if anyone else has been yielding f1 training scores that are lower than their f1 validation scores with this script. The difference is small, but I would expect it to be the other way around??</p>\n<p>Thanks in advance for any help!</p>",
      "rawMarkdown": "Curious if anyone else has been yielding f1 training scores that are lower than their f1 validation scores with this script. The difference is small, but I would expect it to be the other way around??\n\nThanks in advance for any help!",
      "replies": [
        {
          "id": 1313663,
          "postDate": "2021-05-18T16:40:42.197Z",
          "content": "<p>Purely guessing here since I don't know your training pipeline, but it could be that you are applying some augmentation methods(to make the model to overfit less) during training but not during validation. If the discrepancy is small, I wouldn't worry too much about it. I had the same issue during the initial epochs and I think later on your validation score should become lower than your training score.</p>",
          "rawMarkdown": "Purely guessing here since I don't know your training pipeline, but it could be that you are applying some augmentation methods(to make the model to overfit less) during training but not during validation. If the discrepancy is small, I wouldn't worry too much about it. I had the same issue during the initial epochs and I think later on your validation score should become lower than your training score.",
          "votes": 4
        },
        {
          "id": 1313733,
          "postDate": "2021-05-18T17:23:07Z",
          "content": "<p>Thanks Derek!</p>\n<p>Glad to hear that you are getting the same trend. I have so far just played with mixup (\"modified\" as per the implementation in <a href=\"https://www.kaggle.com/theoviel/training-a-winning-model?scriptVersionId=42814701\" target=\"_blank\">Theo's 3rd place solution last year</a>). I am also going to experiment with spec striping augmentations. </p>\n<p>Would there be any reason to apply test time augmentation with either of these methods? It looks like Theo did not do test time augmentation for mixup or striping.</p>\n<p>Thanks again for the help, especially as I am a relative beginner with deep learning.</p>\n<p>Daniel  </p>",
          "rawMarkdown": "Thanks Derek!\n\nGlad to hear that you are getting the same trend. I have so far just played with mixup (\"modified\" as per the implementation in [Theo's 3rd place solution last year](https://www.kaggle.com/theoviel/training-a-winning-model?scriptVersionId=42814701)). I am also going to experiment with spec striping augmentations. \n\nWould there be any reason to apply test time augmentation with either of these methods? It looks like Theo did not do test time augmentation for mixup or striping.\n\nThanks again for the help, especially as I am a relative beginner with deep learning.\n\nDaniel  ",
          "votes": 1
        },
        {
          "id": 1313879,
          "postDate": "2021-05-18T18:58:02Z",
          "content": "<p>I'm going to give my 2 cents(other ppl please correct me if I'm wrong). The reason for applying test time augmentation is to make sure that the augmented test data \"resembles\" the training data, whereas applying training data augmentation is to reduce overfitting during training. In this competition, the distributions between the test data and training data are drastically different due to the quality of recordings. Therefore, it's quite hard to make the test resembles the training data. That's probably the reason why you don't want to apply test time augmentation in this competition.</p>",
          "rawMarkdown": "I'm going to give my 2 cents(other ppl please correct me if I'm wrong). The reason for applying test time augmentation is to make sure that the augmented test data \"resembles\" the training data, whereas applying training data augmentation is to reduce overfitting during training. In this competition, the distributions between the test data and training data are drastically different due to the quality of recordings. Therefore, it's quite hard to make the test resembles the training data. That's probably the reason why you don't want to apply test time augmentation in this competition.",
          "votes": 6
        },
        {
          "id": 1314012,
          "postDate": "2021-05-18T21:45:15.093Z",
          "content": "<p>Makes sense to me Derek! I guess it’s worth noting I was originally referencing the various metrics logged during training, not at the inference stage. For example, I was looking to compare the training vs. validation differences within my local cross validation to asses the level overfitting. Picked that up from a kaggle days talk by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>.</p>",
          "rawMarkdown": "Makes sense to me Derek! I guess it’s worth noting I was originally referencing the various metrics logged during training, not at the inference stage. For example, I was looking to compare the training vs. validation differences within my local cross validation to asses the level overfitting. Picked that up from a kaggle days talk by @cpmpml.",
          "votes": 2
        },
        {
          "id": 1314023,
          "postDate": "2021-05-18T22:03:03.713Z",
          "content": "<p>^ This may be a result of not performing enough epochs at first. When I extended the number of epochs the validation set metrics eventually became lower than the training set metrics, by the ~30th epoch for the F1 statistic. </p>",
          "rawMarkdown": "^ This may be a result of not performing enough epochs at first. When I extended the number of epochs the validation set metrics eventually became lower than the training set metrics, by the ~30th epoch for the F1 statistic. "
        },
        {
          "id": 1314031,
          "postDate": "2021-05-18T22:19:12.800Z",
          "content": "<p>Also it's important to note here that good validation score on the training set doesn't necessarily translate to good score on the test data because these 2 sets have different distributions. In fact, it's possible for a low validation score epoch model weight to outperform a high validation score epoch model weight when used on the test data.</p>",
          "rawMarkdown": "Also it's important to note here that good validation score on the training set doesn't necessarily translate to good score on the test data because these 2 sets have different distributions. In fact, it's possible for a low validation score epoch model weight to outperform a high validation score epoch model weight when used on the test data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1312317,
      "postDate": "2021-05-18T01:46:39.020Z",
      "content": "<p>Thank you for sharing awesome datasets! Because I start this competition now, I am considering whether I should use this dataset.<br>\nMy concern is: will utilizing this dataset hinder us from waveform transforms, though still we can apply spec augmenter ?<br>\nHow do you augment data with this dataset? Or waveform transforms are less useful in this competition?</p>",
      "rawMarkdown": "Thank you for sharing awesome datasets! Because I start this competition now, I am considering whether I should use this dataset.\nMy concern is: will utilizing this dataset hinder us from waveform transforms, though still we can apply spec augmenter ?\nHow do you augment data with this dataset? Or waveform transforms are less useful in this competition?"
    },
    {
      "id": 1300491,
      "postDate": "2021-05-10T13:57:31.283Z",
      "content": "<p>Thanks for your sharing.What I worry about is whether it will lose a lot of details when converting sound to image.It's my first time working with audio😂</p>",
      "rawMarkdown": "Thanks for your sharing.What I worry about is whether it will lose a lot of details when converting sound to image.It's my first time working with audio😂",
      "replies": [
        {
          "id": 1304322,
          "postDate": "2021-05-12T14:54:11.287Z",
          "content": "<p>All the winning solutions of past audio competitions were based on mels specs images, so we can assume with little risk that the information loss is harmless.</p>",
          "rawMarkdown": "All the winning solutions of past audio competitions were based on mels specs images, so we can assume with little risk that the information loss is harmless."
        }
      ]
    },
    {
      "id": 1298747,
      "postDate": "2021-05-09T06:48:16.283Z",
      "content": "<p>Hi! Thank you for the dataset and code! I was wondering if you were able to train the models on colab successfully. I tried your notebook on Kaggle kernel, and it worked well, but with the same code on colab, the performance dropped significantly. I was wondering if you have experienced similar things. Here're some details: <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/237507\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/237507</a> </p>\n<p>Thank you very much!</p>",
      "rawMarkdown": "Hi! Thank you for the dataset and code! I was wondering if you were able to train the models on colab successfully. I tried your notebook on Kaggle kernel, and it worked well, but with the same code on colab, the performance dropped significantly. I was wondering if you have experienced similar things. Here're some details: https://www.kaggle.com/c/birdclef-2021/discussion/237507 \n\nThank you very much!",
      "replies": [
        {
          "id": 1298777,
          "postDate": "2021-05-09T07:32:08.607Z",
          "content": "<p>This is definitely weird … I didn't have this issue (perhaps I have it but I didn't realize)  since I have no model trained on Kaggle. All my models are trained on Colab. Perhaps, try lowering your learning rate a litlle bit could change the learning path. But this is just a speculation.</p>\n<p>Sorry I can't help you more.</p>",
          "rawMarkdown": "This is definitely weird ... I didn't have this issue (perhaps I have it but I didn't realize)  since I have no model trained on Kaggle. All my models are trained on Colab. Perhaps, try lowering your learning rate a litlle bit could change the learning path. But this is just a speculation.\n\nSorry I can't help you more.",
          "votes": 1
        },
        {
          "id": 1298838,
          "postDate": "2021-05-09T08:57:29.423Z",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> Thanks for your reply! It is reasssuring that you are able to create good results with colab. I will continue the investigation then.</p>",
          "rawMarkdown": "@kneroma Thanks for your reply! It is reasssuring that you are able to create good results with colab. I will continue the investigation then."
        }
      ]
    },
    {
      "id": 1283066,
      "postDate": "2021-04-24T14:50:57.500Z",
      "content": "<p>hi <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> !<br>\nthank you very much for sharing!<br>\nhow do you deal with different shapes of .npy files during training?</p>",
      "rawMarkdown": "hi @kneroma !\nthank you very much for sharing!\nhow do you deal with different shapes of .npy files during training?",
      "replies": [
        {
          "id": 1283280,
          "postDate": "2021-04-24T19:01:58.447Z",
          "content": "<p>Each image is from a 7s windows so their shapes are the same. Which different shapes are you talking about ?</p>",
          "rawMarkdown": "Each image is from a 7s windows so their shapes are the same. Which different shapes are you talking about ?",
          "votes": 1
        },
        {
          "id": 1283288,
          "postDate": "2021-04-24T19:22:39.113Z",
          "content": "<p>I mean some file has (13, 128, 281) shape, it means there are 13 7s windows of 1 audio, right?<br>\nother file could have (7, 128, 281) shape, etc<br>\nto train models we need the same shape of input, that's why my question</p>",
          "rawMarkdown": "I mean some file has (13, 128, 281) shape, it means there are 13 7s windows of 1 audio, right?\nother file could have (7, 128, 281) shape, etc\nto train models we need the same shape of input, that's why my question"
        },
        {
          "id": 1283368,
          "postDate": "2021-04-24T21:32:45.627Z",
          "content": "<p>probably these are the no. of recordings per class - how did you derive to this? <br>\nanyway, the melspecs from the dataset output are: <code>(3, 128, 281)</code> 3-channel images 128x281, i.e. with 128 mel filters  (see cell #24 in mentioned notebook) </p>",
          "rawMarkdown": "probably these are the no. of recordings per class - how did you derive to this? \nanyway, the melspecs from the dataset output are: `(3, 128, 281)` 3-channel images 128x281, i.e. with 128 mel filters  (see cell #24 in mentioned notebook) "
        },
        {
          "id": 1287053,
          "postDate": "2021-04-28T16:49:36.027Z",
          "content": "<p>I think the data only has one channel (right?), and we have to concat the arrays with np.vstack((array1, array2)) for example.</p>\n<p>Thank you very much <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> !!</p>",
          "rawMarkdown": "I think the data only has one channel (right?), and we have to concat the arrays with np.vstack((array1, array2)) for example.\n\nThank you very much @kneroma !!"
        },
        {
          "id": 1287062,
          "postDate": "2021-04-28T16:58:05.787Z",
          "content": "<p>Yes : (13, 128, 281) = (<strong>num records</strong>, <strong>freq dim</strong>, <strong>time dim</strong>). We only take a single record (random) at each epoch (<strong>freq dim</strong>, <strong>time dim</strong>), this record is converted into a 3 channels image in the <strong>nomalize</strong> function.</p>",
          "rawMarkdown": "Yes : (13, 128, 281) = (**num records**, **freq dim**, **time dim**). We only take a single record (random) at each epoch (**freq dim**, **time dim**), this record is converted into a 3 channels image in the **nomalize** function.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1283064,
      "postDate": "2021-04-24T14:49:54.997Z",
      "content": "<p>Do you extract spectrograms from audio with a duration of 7 seconds randomly?</p>",
      "rawMarkdown": "Do you extract spectrograms from audio with a duration of 7 seconds randomly?",
      "replies": [
        {
          "id": 1283279,
          "postDate": "2021-04-24T19:00:19.407Z",
          "content": "<p>For short audios there is no randomness involved. For longer ones, I took a 7*10 random window (the starting point is random).</p>",
          "rawMarkdown": "For short audios there is no randomness involved. For longer ones, I took a 7*10 random window (the starting point is random).",
          "votes": 1
        },
        {
          "id": 1283460,
          "postDate": "2021-04-25T00:25:50.190Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1283485,
          "postDate": "2021-04-25T01:15:21.413Z",
          "content": "<p>For a long audio, then how do you label each random chunk of the audio? You know the label is for the whole audio.</p>",
          "rawMarkdown": "For a long audio, then how do you label each random chunk of the audio? You know the label is for the whole audio."
        },
        {
          "id": 1287063,
          "postDate": "2021-04-28T16:59:01.073Z",
          "content": "<p>As you said it, the label is for the whole audio, thus is for the chunk as well :) </p>",
          "rawMarkdown": "As you said it, the label is for the whole audio, thus is for the chunk as well :) ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1283010,
      "postDate": "2021-04-24T13:50:22.280Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1285218,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2021-04-26T17:40:29.590000",
      "content": "<p>In case anyone's using tensorflow datasets:</p>\n<p>I recently (re-)ran some performance experiments with a desktop GPU, and found it fastest to apply time-domain augmentations in the TF dataset (which runs with heavily parallelization on the machine's CPU), and then apply the STFT+MelSpec extraction in the model code (which runs on the GPU: The STFT+Melspec ops are basically just batched convolution and matmul, which is exactly whet the GPU is best at). </p>\n<p>Producing a batch of 32 time-domain augmented audio segments now takes ~0.08 seconds. When the melspec operations are included in the dataset it takes about 0.4 seconds to produce a batch.  The time domain augmentations include mix-up addition of two examples, gain randomization, and noise addition. I've also got a few melspec-domain augmentations, including emulating a random low-pass filter in the mel domain, which happen in the model graph now.</p>\n<p>Meanwhile, once the melspec computation is in the model graph, the model trains at ~0.23s per batch, so isn't input bound. Obvs, all the perf numbers will depend on your particular machine and model architecture; the point is that if you can keep the model well-fed with time-domain data, it may be faster to make the melspec part of the model so it runs on GPU/TPU.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1290385,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-01T22:18:11.227000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> for your suggestions. For pytorch users who want to compute the mels on GPU, there are <a href=\"https://pytorch.org/audio/stable/index.html\" target=\"_blank\">torchaudio</a> and <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">torch-audiomentations</a> .</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1282331,
      "author_name": "hide on bread",
      "author_url": "",
      "post_date": "2021-04-23T19:44:45.030000",
      "content": "<p>Thanks for sharing this technique! I'm wondering if converting mels to np.uint8 would result in the loss of information, which would then negatively impact the model. Is this a trade off between training time and model accuracy?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1282466,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-24T01:11:17.337000",
          "content": "<p>A good point over there. But, for me, even without the pre-computation, I was using cv2 images so it does not change my pipeline. I won't expect the information loss to be too much damaging.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1311108,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2021-05-17T06:53:33.883000",
      "content": "<p>Hi, your notebook is great for novice like me. I have a small question, in <a href=\"https://www.kaggle.com/kneroma/birdclef-mels-computer-public\" target=\"_blank\">BirdCLEF Mels Computer [Public]\n</a> you call <code>image = mono_to_color(melspec)</code> without mean and std. </p>\n<pre><code>def mono_to_color(X, eps=1e-6, mean=None, std=None):\n    mean = mean or X.mean()\n    std = std or X.std()\n    X = (X - mean) / (std + eps)\n\n    _min, _max = X.min(), X.max()\n\n    if (_max - _min) &gt; eps:\n        V = np.clip(X, _min, _max)\n        V = 255 * (V - _min) / (_max - _min)\n        V = V.astype(np.uint8)\n    else:\n        V = np.zeros_like(X, dtype=np.uint8)\n\n    return V\n</code></pre>\n<p>This is the first time for me to handle audio data, I find some <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/177551\" target=\"_blank\">discussion</a> that normalize in another way. It it ok to normalize with self's mean and std?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1312257,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-18T00:15:22.847000",
          "content": "<p>Great question.</p>\n<p>I've also seen theo's discussion. Personally I'm not fan of carrying ImageNet's mean and std everywhere for normalization purposes. Here, we're not doing a simple transfer learning (the model will be trained on a completely different dataset for many and many epochs) and the impact of the pretrained weights is questionable. </p>\n<p>I would say It's much a question of taste. I would be happy to read someone's results after moving from the original <strong>mono_to_color</strong> to the new one.</p>\n<p>The fact that the means and stds are not constant over the samples is actually a feature and not a bug as the mels are not true images. A mel is semantically the same when divided by a constant or not, same for addition. Some people are doing even more complicated self transformations, like raising the whole mels into the power of 3.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1310767,
      "author_name": "Daniel Furman",
      "author_url": "",
      "post_date": "2021-05-16T23:51:52.803000",
      "content": "<p>I \"ran all\" and got an error at the last step. Seemed like resnest wasn't working, because I then tried a resnet50 and that is working fine. Any ideas why the last code bubble would throw an error?</p>\n<p>Thanks in advance for the help !</p>\n<p>The error specifically was that the Resnest50 downloader only got to 5/6 and then failed, yet the exception error was not thrown for some reason</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1312266,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-18T00:27:29.853000",
          "content": "<p>I think the resnest50 pretrained weights' link is broken.</p>\n<p>Please refer to <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/comments#1303631\" target=\"_blank\">this thread</a> for a workaround.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1305644,
      "author_name": "Nin7a1",
      "author_url": "",
      "post_date": "2021-05-13T12:06:18.607000",
      "content": "<p>Thanks for your sharing! I also have a question, are data augmentations like Pink noise/Gaussian noise/Gaussian SNR/… applied before mels being computed?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1313111,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-18T12:10:31.937000",
          "content": "<p>No. Still you can use spec augs. If you're interested in raw audio augs, precomputed mels are not a good choice.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1290016,
      "author_name": "Gurgel",
      "author_url": "",
      "post_date": "2021-05-01T14:56:23.903000",
      "content": "<p>Thank you for sharing all this, I am learning a lot from you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1290386,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-01T22:18:56.080000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/daniellga\" target=\"_blank\">@daniellga</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1289816,
      "author_name": "Moein",
      "author_url": "",
      "post_date": "2021-05-01T11:54:00.690000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>, thanks for sharing these valuable resources! I had a few questions. I've joined this competition recently and it's my first time working with audio, so sorry in advance if my questions are stupid :)</p>\n<p>1- Doesn't it harm the model when it gets trained on random 7 seconds of the recording where we do not know for sure that if there is any bird present? I mean the randomly selected 7 sec can be from a part of the recording with pure noise and the model sees a real bird label for that noise. Of course some of these 7 seconds do have the bird but what about those that do not? And if I'm correct in the way I see the problem, are there alternatives to your strategy?</p>\n<p>2- In the notebook that you shared which builds these datasets, in the last function you are setting the <code>step</code> variable to DURATION * 0.666 * SR. What is that 0.666 about? Does it mean that the final separated records will be less than 7 secs?</p>\n<p>3- For this last one I do not really expect an answer and answer this if you're okay with sharing: Is your LB (or near that) has been achieved using these datasets?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1290379,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-01T22:15:05.683000",
          "content": "<ol>\n<li><p>The main problem is that the train set is weakly labeled (target is only available for the whole clip, not for audio segments). Random  chunk is a very basic startegy, and I won't expect it to be the smartest one. There may be several alternatives. For instance, you may train a <strong>call / no call</strong> classifier and afterward consider only the chunks with a call …</p></li>\n<li><p>0.666 is meaningless, I choose this value just to avoid a memory overflow on Kaggle. Generally, if <strong>step &lt; DURATION*SR</strong>, it means that there is an overlapping between your chunks. Depending on how much memory do you have, you can lower this 0.666 to get more overlapped chunks.</p></li>\n<li><p>My first subs was exclusively based on those datasets (I got 0.73 LB with them). Currently, I'm experimenting many ideas locally, with no great succes xD .</p></li>\n</ol>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1290525,
          "author_name": "Moein",
          "author_url": "",
          "post_date": "2021-05-02T05:00:59.540000",
          "content": "<p>Thank you very much!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1291361,
          "author_name": "Steve Pousty",
          "author_url": "",
          "post_date": "2021-05-03T01:15:10.237000",
          "content": "<p>I tried using an onset analyzer to break it apart but it turned out breaking up the chickadee call into small little segments. <br>\nCode examples here:<br>\n<a href=\"https://github.com/thesteve0/birdclef21/commit/8f7e0e129243bfdcc1f675c007ccd837b60aaa24?branch=8f7e0e129243bfdcc1f675c007ccd837b60aaa24&amp;diff=unified\" target=\"_blank\">https://github.com/thesteve0/birdclef21/commit/8f7e0e129243bfdcc1f675c007ccd837b60aaa24?branch=8f7e0e129243bfdcc1f675c007ccd837b60aaa24&amp;diff=unified</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307648,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-05-14T15:24:41.260000",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  how much was ur local cv that gave lb 0.73 , is that five fold score or single fold thanks in advance</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1283391,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2021-04-24T22:07:57.213000",
      "content": "<p>Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1283465,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-25T00:29:01.730000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1281847,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-04-23T11:10:12.190000",
      "content": "<p>Thank for sharing <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> - that's some kinda of <code>kkiller-guns</code> :) very generous of you! <br>\nI wish you'll get what you deserved at the end of the competition</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1282426,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-23T23:15:40.773000",
          "content": "<p>Thanks for the kind words <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1281431,
      "author_name": "Adriel Vieira",
      "author_url": "",
      "post_date": "2021-04-22T23:50:57.827000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>! That's very useful =)</p>\n<p>It looks like the \"7 seconds records numpy images with truncation\" link is broken, though</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1281438,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-23T00:03:55.503000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/adriel\" target=\"_blank\">@adriel</a> . Link fixed !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1304809,
      "author_name": "tiagovieira",
      "author_url": "",
      "post_date": "2021-05-12T22:36:28.690000",
      "content": "<p>This is great, I have a question, is there any advantage if we convert the mel spectrogram directly to a tensor?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1304988,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-13T03:49:07.587000",
          "content": "<p>The only advantage I can see is that it could avoid the numpy - torch conversion. But since pytorch tensor and numpy arrays can share the same storage, the overhead brings by the conversion step may be negligible.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1306230,
          "author_name": "tiagovieira",
          "author_url": "",
          "post_date": "2021-05-13T17:15:53.843000",
          "content": "<p>Quick question, is it possible to use the preprocessed dataset in npy on the submission? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1313115,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-18T12:12:26.080000",
          "content": "<p>We don't have access to the hidden test set, except inside the \"submission\" environment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304286,
      "author_name": "majunfu",
      "author_url": "",
      "post_date": "2021-05-12T14:34:12.573000",
      "content": "<p>I am learning a lot from you too! Thanks for your sharing.It's very useful for me.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1304319,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-12T14:52:34.347000",
          "content": "<p>Thanks  <a href=\"https://www.kaggle.com/majunfu\" target=\"_blank\">@majunfu</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1305695,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-05-13T12:34:46.320000",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  thanku for resources.<br>\nOne ask. <br>\nDuring inference i see there 20 test simulating ogg files broken into 120 samples each lasting 5 sec.<br>\nDoes a  each datafile broken into 120 samples is either no call or call  from a single bird </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307580,
          "author_name": "majunfu",
          "author_url": "",
          "post_date": "2021-05-14T14:34:23.073000",
          "content": "<p>Hi,I want to know if your model is also obtained through current image training?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1318909,
      "author_name": "liuze",
      "author_url": "",
      "post_date": "2021-05-22T17:03:24.760000",
      "content": "<p>thanks for sharing, one question, what's the difference between truncation and no truncation?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1313614,
      "author_name": "Daniel Furman",
      "author_url": "",
      "post_date": "2021-05-18T16:14:54.223000",
      "content": "<p>Curious if anyone else has been yielding f1 training scores that are lower than their f1 validation scores with this script. The difference is small, but I would expect it to be the other way around??</p>\n<p>Thanks in advance for any help!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1313663,
          "author_name": "hide on bread",
          "author_url": "",
          "post_date": "2021-05-18T16:40:42.197000",
          "content": "<p>Purely guessing here since I don't know your training pipeline, but it could be that you are applying some augmentation methods(to make the model to overfit less) during training but not during validation. If the discrepancy is small, I wouldn't worry too much about it. I had the same issue during the initial epochs and I think later on your validation score should become lower than your training score.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1313733,
          "author_name": "Daniel Furman",
          "author_url": "",
          "post_date": "2021-05-18T17:23:07",
          "content": "<p>Thanks Derek!</p>\n<p>Glad to hear that you are getting the same trend. I have so far just played with mixup (\"modified\" as per the implementation in <a href=\"https://www.kaggle.com/theoviel/training-a-winning-model?scriptVersionId=42814701\" target=\"_blank\">Theo's 3rd place solution last year</a>). I am also going to experiment with spec striping augmentations. </p>\n<p>Would there be any reason to apply test time augmentation with either of these methods? It looks like Theo did not do test time augmentation for mixup or striping.</p>\n<p>Thanks again for the help, especially as I am a relative beginner with deep learning.</p>\n<p>Daniel  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1313879,
          "author_name": "hide on bread",
          "author_url": "",
          "post_date": "2021-05-18T18:58:02",
          "content": "<p>I'm going to give my 2 cents(other ppl please correct me if I'm wrong). The reason for applying test time augmentation is to make sure that the augmented test data \"resembles\" the training data, whereas applying training data augmentation is to reduce overfitting during training. In this competition, the distributions between the test data and training data are drastically different due to the quality of recordings. Therefore, it's quite hard to make the test resembles the training data. That's probably the reason why you don't want to apply test time augmentation in this competition.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1314012,
          "author_name": "Daniel Furman",
          "author_url": "",
          "post_date": "2021-05-18T21:45:15.093000",
          "content": "<p>Makes sense to me Derek! I guess it’s worth noting I was originally referencing the various metrics logged during training, not at the inference stage. For example, I was looking to compare the training vs. validation differences within my local cross validation to asses the level overfitting. Picked that up from a kaggle days talk by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1314023,
          "author_name": "Daniel Furman",
          "author_url": "",
          "post_date": "2021-05-18T22:03:03.713000",
          "content": "<p>^ This may be a result of not performing enough epochs at first. When I extended the number of epochs the validation set metrics eventually became lower than the training set metrics, by the ~30th epoch for the F1 statistic. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1314031,
          "author_name": "hide on bread",
          "author_url": "",
          "post_date": "2021-05-18T22:19:12.800000",
          "content": "<p>Also it's important to note here that good validation score on the training set doesn't necessarily translate to good score on the test data because these 2 sets have different distributions. In fact, it's possible for a low validation score epoch model weight to outperform a high validation score epoch model weight when used on the test data.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1312317,
      "author_name": "Yoheiii",
      "author_url": "",
      "post_date": "2021-05-18T01:46:39.020000",
      "content": "<p>Thank you for sharing awesome datasets! Because I start this competition now, I am considering whether I should use this dataset.<br>\nMy concern is: will utilizing this dataset hinder us from waveform transforms, though still we can apply spec augmenter ?<br>\nHow do you augment data with this dataset? Or waveform transforms are less useful in this competition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1300491,
      "author_name": "majunfu",
      "author_url": "",
      "post_date": "2021-05-10T13:57:31.283000",
      "content": "<p>Thanks for your sharing.What I worry about is whether it will lose a lot of details when converting sound to image.It's my first time working with audio😂</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1304322,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-12T14:54:11.287000",
          "content": "<p>All the winning solutions of past audio competitions were based on mels specs images, so we can assume with little risk that the information loss is harmless.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1298747,
      "author_name": "Gold Retriever",
      "author_url": "",
      "post_date": "2021-05-09T06:48:16.283000",
      "content": "<p>Hi! Thank you for the dataset and code! I was wondering if you were able to train the models on colab successfully. I tried your notebook on Kaggle kernel, and it worked well, but with the same code on colab, the performance dropped significantly. I was wondering if you have experienced similar things. Here're some details: <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/237507\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/237507</a> </p>\n<p>Thank you very much!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1298777,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-05-09T07:32:08.607000",
          "content": "<p>This is definitely weird … I didn't have this issue (perhaps I have it but I didn't realize)  since I have no model trained on Kaggle. All my models are trained on Colab. Perhaps, try lowering your learning rate a litlle bit could change the learning path. But this is just a speculation.</p>\n<p>Sorry I can't help you more.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298838,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2021-05-09T08:57:29.423000",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> Thanks for your reply! It is reasssuring that you are able to create good results with colab. I will continue the investigation then.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1283066,
      "author_name": "Mikhail Pchelintsev",
      "author_url": "",
      "post_date": "2021-04-24T14:50:57.500000",
      "content": "<p>hi <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> !<br>\nthank you very much for sharing!<br>\nhow do you deal with different shapes of .npy files during training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1283280,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-24T19:01:58.447000",
          "content": "<p>Each image is from a 7s windows so their shapes are the same. Which different shapes are you talking about ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1283288,
          "author_name": "Mikhail Pchelintsev",
          "author_url": "",
          "post_date": "2021-04-24T19:22:39.113000",
          "content": "<p>I mean some file has (13, 128, 281) shape, it means there are 13 7s windows of 1 audio, right?<br>\nother file could have (7, 128, 281) shape, etc<br>\nto train models we need the same shape of input, that's why my question</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1283368,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2021-04-24T21:32:45.627000",
          "content": "<p>probably these are the no. of recordings per class - how did you derive to this? <br>\nanyway, the melspecs from the dataset output are: <code>(3, 128, 281)</code> 3-channel images 128x281, i.e. with 128 mel filters  (see cell #24 in mentioned notebook) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287053,
          "author_name": "Yamil Honaine Bórquez",
          "author_url": "",
          "post_date": "2021-04-28T16:49:36.027000",
          "content": "<p>I think the data only has one channel (right?), and we have to concat the arrays with np.vstack((array1, array2)) for example.</p>\n<p>Thank you very much <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> !!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287062,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-28T16:58:05.787000",
          "content": "<p>Yes : (13, 128, 281) = (<strong>num records</strong>, <strong>freq dim</strong>, <strong>time dim</strong>). We only take a single record (random) at each epoch (<strong>freq dim</strong>, <strong>time dim</strong>), this record is converted into a 3 channels image in the <strong>nomalize</strong> function.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1283064,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2021-04-24T14:49:54.997000",
      "content": "<p>Do you extract spectrograms from audio with a duration of 7 seconds randomly?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1283279,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-24T19:00:19.407000",
          "content": "<p>For short audios there is no randomness involved. For longer ones, I took a 7*10 random window (the starting point is random).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1283460,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-25T00:25:50.190000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1283485,
          "author_name": "XiaokangWang",
          "author_url": "",
          "post_date": "2021-04-25T01:15:21.413000",
          "content": "<p>For a long audio, then how do you label each random chunk of the audio? You know the label is for the whole audio.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287063,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2021-04-28T16:59:01.073000",
          "content": "<p>As you said it, the label is for the whole audio, thus is for the chunk as well :) </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1283010,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-24T13:50:22.280000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1281413": "My training pipeline was used to last 9 hours for just 12 epochs. This duration has been incredibly reduced when I precomptued the mels and converted them into \"**np.uint8**\" dtypes (one of the smallest  numpy dtypes). Now I can train a [ full model on Kaggle](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab) and one fold of 20 epochs could last **less than two hours**.\n\nHere are the links to those handy datasets:\n\n* [7 seconds records numpy images with truncation](https://www.kaggle.com/kneroma/kkiller-birdclef-2021)\n\nIf you're interrested in the whole record melspecs images (no truncation):\n\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part1\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part2\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part3\nhttps://www.kaggle.com/kneroma/kkiller-birdclef-mels-computer-d7-part4\n",
    "1285218": "In case anyone's using tensorflow datasets:\n\nI recently (re-)ran some performance experiments with a desktop GPU, and found it fastest to apply time-domain augmentations in the TF dataset (which runs with heavily parallelization on the machine's CPU), and then apply the STFT+MelSpec extraction in the model code (which runs on the GPU: The STFT+Melspec ops are basically just batched convolution and matmul, which is exactly whet the GPU is best at). \n\nProducing a batch of 32 time-domain augmented audio segments now takes ~0.08 seconds. When the melspec operations are included in the dataset it takes about 0.4 seconds to produce a batch.  The time domain augmentations include mix-up addition of two examples, gain randomization, and noise addition. I've also got a few melspec-domain augmentations, including emulating a random low-pass filter in the mel domain, which happen in the model graph now.\n\nMeanwhile, once the melspec computation is in the model graph, the model trains at ~0.23s per batch, so isn't input bound. Obvs, all the perf numbers will depend on your particular machine and model architecture; the point is that if you can keep the model well-fed with time-domain data, it may be faster to make the melspec part of the model so it runs on GPU/TPU.",
    "1282331": "Thanks for sharing this technique! I'm wondering if converting mels to np.uint8 would result in the loss of information, which would then negatively impact the model. Is this a trade off between training time and model accuracy?",
    "1311108": "Hi, your notebook is great for novice like me. I have a small question, in [BirdCLEF Mels Computer [Public]\n](https://www.kaggle.com/kneroma/birdclef-mels-computer-public) you call `image = mono_to_color(melspec)` without mean and std. \n```\ndef mono_to_color(X, eps=1e-6, mean=None, std=None):\n    mean = mean or X.mean()\n    std = std or X.std()\n    X = (X - mean) / (std + eps)\n    \n    _min, _max = X.min(), X.max()\n\n    if (_max - _min) > eps:\n        V = np.clip(X, _min, _max)\n        V = 255 * (V - _min) / (_max - _min)\n        V = V.astype(np.uint8)\n    else:\n        V = np.zeros_like(X, dtype=np.uint8)\n\n    return V\n```\nThis is the first time for me to handle audio data, I find some [discussion](https://www.kaggle.com/c/birdsong-recognition/discussion/177551) that normalize in another way. It it ok to normalize with self's mean and std?",
    "1310767": "I \"ran all\" and got an error at the last step. Seemed like resnest wasn't working, because I then tried a resnet50 and that is working fine. Any ideas why the last code bubble would throw an error?\n\nThanks in advance for the help !\n\nThe error specifically was that the Resnest50 downloader only got to 5/6 and then failed, yet the exception error was not thrown for some reason\n\n\n",
    "1305644": "Thanks for your sharing! I also have a question, are data augmentations like Pink noise/Gaussian noise/Gaussian SNR/... applied before mels being computed?",
    "1290016": "Thank you for sharing all this, I am learning a lot from you!",
    "1289816": "Hey @kneroma, thanks for sharing these valuable resources! I had a few questions. I've joined this competition recently and it's my first time working with audio, so sorry in advance if my questions are stupid :)\n\n1- Doesn't it harm the model when it gets trained on random 7 seconds of the recording where we do not know for sure that if there is any bird present? I mean the randomly selected 7 sec can be from a part of the recording with pure noise and the model sees a real bird label for that noise. Of course some of these 7 seconds do have the bird but what about those that do not? And if I'm correct in the way I see the problem, are there alternatives to your strategy?\n\n2- In the notebook that you shared which builds these datasets, in the last function you are setting the `step` variable to DURATION * 0.666 * SR. What is that 0.666 about? Does it mean that the final separated records will be less than 7 secs?\n\n3- For this last one I do not really expect an answer and answer this if you're okay with sharing: Is your LB (or near that) has been achieved using these datasets?",
    "1283391": "Thank you!",
    "1281847": "Thank for sharing @kneroma - that's some kinda of `kkiller-guns` :) very generous of you! \nI wish you'll get what you deserved at the end of the competition",
    "1281431": "Thanks @kneroma! That's very useful =)\n\nIt looks like the \"7 seconds records numpy images with truncation\" link is broken, though",
    "1304809": "This is great, I have a question, is there any advantage if we convert the mel spectrogram directly to a tensor?",
    "1304286": "I am learning a lot from you too! Thanks for your sharing.It's very useful for me.",
    "1318909": "thanks for sharing, one question, what's the difference between truncation and no truncation?",
    "1313614": "Curious if anyone else has been yielding f1 training scores that are lower than their f1 validation scores with this script. The difference is small, but I would expect it to be the other way around??\n\nThanks in advance for any help!",
    "1312317": "Thank you for sharing awesome datasets! Because I start this competition now, I am considering whether I should use this dataset.\nMy concern is: will utilizing this dataset hinder us from waveform transforms, though still we can apply spec augmenter ?\nHow do you augment data with this dataset? Or waveform transforms are less useful in this competition?",
    "1300491": "Thanks for your sharing.What I worry about is whether it will lose a lot of details when converting sound to image.It's my first time working with audio😂",
    "1298747": "Hi! Thank you for the dataset and code! I was wondering if you were able to train the models on colab successfully. I tried your notebook on Kaggle kernel, and it worked well, but with the same code on colab, the performance dropped significantly. I was wondering if you have experienced similar things. Here're some details: https://www.kaggle.com/c/birdclef-2021/discussion/237507 \n\nThank you very much!",
    "1283066": "hi @kneroma !\nthank you very much for sharing!\nhow do you deal with different shapes of .npy files during training?",
    "1283064": "Do you extract spectrograms from audio with a duration of 7 seconds randomly?",
    "1283010": ""
  }
}