{
  "id": 209206,
  "title": "Precomputing audio files vs Data Augmentation on the fly",
  "url": "/competitions/rfcx-species-audio-detection/discussion/209206",
  "author_name": "",
  "post_date": "2021-01-06T17:12:00.833583500Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I was precomputing the audio files and save them to the file. Then, I was loading them in the dataloader. But I have added a data augmentation to my pipeline and because of that I can not do precompute &amp; load method since I have to apply data augmentation in the process itself. Before the data augmentation my single epoch was taking a minute or so. But now with data augmentation it takes 6 to 7 minutes. </p>\n<p>First of all, is my approach wrong? <br>\nSecondly, what can I do to speed up this process?</p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1141385",
      "postDate": "01/06/2021 17:12:00",
      "content": "<p>Hi everyone,</p>\n<p>I was precomputing the audio files and save them to the file. Then, I was loading them in the dataloader. But I have added a data augmentation to my pipeline and because of that I can not do precompute &amp; load method since I have to apply data augmentation in the process itself. Before the data augmentation my single epoch was taking a minute or so. But now with data augmentation it takes 6 to 7 minutes. </p>\n<p>First of all, is my approach wrong? <br>\nSecondly, what can I do to speed up this process?</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi everyone,\n\nI was precomputing the audio files and save them to the file. Then, I was loading them in the dataloader. But I have added a data augmentation to my pipeline and because of that I can not do precompute & load method since I have to apply data augmentation in the process itself. Before the data augmentation my single epoch was taking a minute or so. But now with data augmentation it takes 6 to 7 minutes. \n\nFirst of all, is my approach wrong? \nSecondly, what can I do to speed up this process?\n\nThanks.",
      "votes": null
    },
    {
      "id": "1142703",
      "postDate": "01/07/2021 14:52:06",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> , good question! I was also thinking about this problem. Did some experiments, and hopefully a bit helpful.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2Fc01dd3b559520a7c99a061806df8a673%2FScreenshot%20from%202021-01-07%2008-39-42.png?generation=1610030415862037&amp;alt=media\" alt=\"\"></p>\n<p>The above notebook demonstrates couple of things,</p>\n<ul>\n<li>The melspectrogram computation is mainly three parts: STFT, absolute&amp;power, convert to Mel scale, and optionally covnert to db (which is taking a log).</li>\n<li>the most time consuming STFT is a linear operation, in the sense that, if we want to add noise to the original input, it's equivalent to add the STFT of noise to the STFT of the original input. So, if noise are from a dictionary of audio clips, then, we can precompute STFT of original inputs and noises, the STFT of their random combination does not needs to be recomputed.</li>\n<li>The absolute&amp;power step is also slow using the original implementation. But if we keep the power to be 2, then, the provided numba implementation is much faster than the numpy one, with minimum discrepancy. I learnt this from <a href=\"https://stackoverflow.com/questions/30437947/most-memory-efficient-way-to-compute-abs2-of-complex-numpy-ndarray\" target=\"_blank\">stackoverflow</a> and there are couple of more implementations.</li>\n</ul>",
      "rawMarkdown": "Hi, @snnclsr , good question! I was also thinking about this problem. Did some experiments, and hopefully a bit helpful.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2Fc01dd3b559520a7c99a061806df8a673%2FScreenshot%20from%202021-01-07%2008-39-42.png?generation=1610030415862037&alt=media)\n\nThe above notebook demonstrates couple of things,\n- The melspectrogram computation is mainly three parts: STFT, absolute&power, convert to Mel scale, and optionally covnert to db (which is taking a log).\n- the most time consuming STFT is a linear operation, in the sense that, if we want to add noise to the original input, it's equivalent to add the STFT of noise to the STFT of the original input. So, if noise are from a dictionary of audio clips, then, we can precompute STFT of original inputs and noises, the STFT of their random combination does not needs to be recomputed.\n- The absolute&power step is also slow using the original implementation. But if we keep the power to be 2, then, the provided numba implementation is much faster than the numpy one, with minimum discrepancy. I learnt this from [stackoverflow](https://stackoverflow.com/questions/30437947/most-memory-efficient-way-to-compute-abs2-of-complex-numpy-ndarray) and there are couple of more implementations.",
      "votes": null
    },
    {
      "id": "1146510",
      "postDate": "01/09/2021 20:10:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> , Thanks a lot for your detailed and great answer! Your answer lighten some dark places in my mind. So instead of using the librosa's melspectrogram function, can we use your version ? I will definitely try btw. </p>\n<p>Also I don't have any prior experience in audio processing. So everything is just new to me :)<br>\nThanks a lot again.</p>",
      "rawMarkdown": "Hi @barnwellguy , Thanks a lot for your detailed and great answer! Your answer lighten some dark places in my mind. So instead of using the librosa's melspectrogram function, can we use your version ? I will definitely try btw. \n\nAlso I don't have any prior experience in audio processing. So everything is just new to me :)\nThanks a lot again.",
      "votes": null
    },
    {
      "id": "1146556",
      "postDate": "01/09/2021 21:15:05",
      "content": "<p>I am glad that you liked it. Feel free to use it if it could be of a little help! It should produce almost identical output as librosa according to my simple tests. I just collected things from internet too :)</p>\n<p>BTW, for some reason, the last <code>power_to_db</code> step can also be faster.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F0c5522e414a0bdc294c663fee9563f98%2FScreenshot%20from%202021-01-09%2015-12-13.png?generation=1610226880419856&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I am glad that you liked it. Feel free to use it if it could be of a little help! It should produce almost identical output as librosa according to my simple tests. I just collected things from internet too :)\n\nBTW, for some reason, the last `power_to_db` step can also be faster.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F0c5522e414a0bdc294c663fee9563f98%2FScreenshot%20from%202021-01-09%2015-12-13.png?generation=1610226880419856&alt=media)",
      "votes": null
    },
    {
      "id": "1146608",
      "postDate": "01/09/2021 21:59:39",
      "content": "<p>Wow, that's great. Understanding/knowing the background of this computations definitely makes everything much more smooth. </p>\n<p>Also, I'm trying to adapt your code to my pipeline. In terms of augmentation, let's say I want to inject some noise to signal. </p>\n<p>As I was searching about this topic, I came to an idea that we can add noise in two ways:</p>\n<ol>\n<li>Adding noise to raw signal like just after reading the audio file.</li>\n<li>Adding noise to the features like melspectrogram (I did not convince myself about this one yet)</li>\n</ol>\n<p>Where should I apply this noise? I will use your variable names here.</p>\n<p>A. wav<br>\nB. fft_windows<br>\nC. s2</p>\n<p>Thanks a lot again 🙏</p>",
      "rawMarkdown": "Wow, that's great. Understanding/knowing the background of this computations definitely makes everything much more smooth. \n\nAlso, I'm trying to adapt your code to my pipeline. In terms of augmentation, let's say I want to inject some noise to signal. \n\nAs I was searching about this topic, I came to an idea that we can add noise in two ways:\n\n1. Adding noise to raw signal like just after reading the audio file.\n2. Adding noise to the features like melspectrogram (I did not convince myself about this one yet)\n\nWhere should I apply this noise? I will use your variable names here.\n\nA. wav\nB. fft_windows\nC. s2\n\nThanks a lot again 🙏",
      "votes": null
    },
    {
      "id": "1146691",
      "postDate": "01/10/2021 00:25:48",
      "content": "<p>I didn't augment the data nearly at all so far in this competition, so I don't have much practical advice. </p>\n<p>But, if we want to go approach 1, which perhaps is most natural, then, of course, injecting noise to wave and do all steps of computation will work. If that is too slow, and we want exactly the same result as going through all steps, then, we can select a collection of typical noises that contains not any target bird call in it and prercompute STFT/fft_windows for it. When we want to inject any of this noise into any sample, we just add the STFT of the noise to the STFT of the sample. </p>\n<p>Since STFT is linear, STFT(sample + noise) = STFT(sample) + STFT(noise)</p>\n<p>So, after precompuation of STFT(sample) and STFT(noise), we don't need to do the most expensive STFT for every sample with noise when we generate data on the go.</p>\n<p>About approach 2, adding noise to <code>s2</code>, or mix sample's <code>s2</code> with other sample's <code>s2</code> is like those typical \"mix-up\" method image processing project does, right? Although, the generated images may not have any corresponding realistic sound which is perhaps not ideal, but it still helps the model to avoid learning trivial spurious characteristics, right? e.g., it is species 5 if the 3rd pixel in the bottom line is between 5.5 and 6.5. Especially, we only have 50 positive data for each species.</p>\n<p>Above are just my 2 cents, without much consideration of this specific context 🤝</p>\n<p>We can discuss more.</p>",
      "rawMarkdown": "I didn't augment the data nearly at all so far in this competition, so I don't have much practical advice. \n\nBut, if we want to go approach 1, which perhaps is most natural, then, of course, injecting noise to wave and do all steps of computation will work. If that is too slow, and we want exactly the same result as going through all steps, then, we can select a collection of typical noises that contains not any target bird call in it and prercompute STFT/fft_windows for it. When we want to inject any of this noise into any sample, we just add the STFT of the noise to the STFT of the sample. \n\nSince STFT is linear, STFT(sample + noise) = STFT(sample) + STFT(noise)\n\nSo, after precompuation of STFT(sample) and STFT(noise), we don't need to do the most expensive STFT for every sample with noise when we generate data on the go.\n\nAbout approach 2, adding noise to `s2`, or mix sample's `s2` with other sample's `s2` is like those typical \"mix-up\" method image processing project does, right? Although, the generated images may not have any corresponding realistic sound which is perhaps not ideal, but it still helps the model to avoid learning trivial spurious characteristics, right? e.g., it is species 5 if the 3rd pixel in the bottom line is between 5.5 and 6.5. Especially, we only have 50 positive data for each species.\n\nAbove are just my 2 cents, without much consideration of this specific context 🤝\n\nWe can discuss more.",
      "votes": null
    },
    {
      "id": "1147980",
      "postDate": "01/10/2021 20:37:32",
      "content": "<p>Hi again, </p>\n<p>First of all, I just want to say thank you. In this phase of the competition, I finally managed to pass 87.6 (the highest scoring public notebook) with your support.<br>\nI went with the option 1 and applied noise to raw signal.  Everything so far seems to work well in my current setup. Things started to make much more sense.</p>\n<p>Thanks a lot again 🙏</p>",
      "rawMarkdown": "Hi again, \n\nFirst of all, I just want to say thank you. In this phase of the competition, I finally managed to pass 87.6 (the highest scoring public notebook) with your support.\nI went with the option 1 and applied noise to raw signal.  Everything so far seems to work well in my current setup. Things started to make much more sense.\n\nThanks a lot again 🙏",
      "votes": null
    },
    {
      "id": "1148205",
      "postDate": "01/11/2021 02:01:30",
      "content": "<p>Congratulations! Happy to hear it helped!<br>\nIt's not easy to beat ensemble submissions! </p>",
      "rawMarkdown": "Congratulations! Happy to hear it helped!\nIt's not easy to beat ensemble submissions!",
      "votes": null
    },
    {
      "id": "1149337",
      "postDate": "01/11/2021 19:15:10",
      "content": "<p>Thanks a lot 😊 </p>\n<p>My solution is also based on ensembles, but there are some improvements on the base models thanks to your ideas.</p>",
      "rawMarkdown": "Thanks a lot 😊 \n\nMy solution is also based on ensembles, but there are some improvements on the base models thanks to your ideas.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1142703,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "01/07/2021 14:52:06",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> , good question! I was also thinking about this problem. Did some experiments, and hopefully a bit helpful.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2Fc01dd3b559520a7c99a061806df8a673%2FScreenshot%20from%202021-01-07%2008-39-42.png?generation=1610030415862037&amp;alt=media\" alt=\"\"></p>\n<p>The above notebook demonstrates couple of things,</p>\n<ul>\n<li>The melspectrogram computation is mainly three parts: STFT, absolute&amp;power, convert to Mel scale, and optionally covnert to db (which is taking a log).</li>\n<li>the most time consuming STFT is a linear operation, in the sense that, if we want to add noise to the original input, it's equivalent to add the STFT of noise to the STFT of the original input. So, if noise are from a dictionary of audio clips, then, we can precompute STFT of original inputs and noises, the STFT of their random combination does not needs to be recomputed.</li>\n<li>The absolute&amp;power step is also slow using the original implementation. But if we keep the power to be 2, then, the provided numba implementation is much faster than the numpy one, with minimum discrepancy. I learnt this from <a href=\"https://stackoverflow.com/questions/30437947/most-memory-efficient-way-to-compute-abs2-of-complex-numpy-ndarray\" target=\"_blank\">stackoverflow</a> and there are couple of more implementations.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1146510,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/09/2021 20:10:36",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> , Thanks a lot for your detailed and great answer! Your answer lighten some dark places in my mind. So instead of using the librosa's melspectrogram function, can we use your version ? I will definitely try btw. </p>\n<p>Also I don't have any prior experience in audio processing. So everything is just new to me :)<br>\nThanks a lot again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146556,
          "author_name": "barnwellguy",
          "author_url": "",
          "post_date": "01/09/2021 21:15:05",
          "content": "<p>I am glad that you liked it. Feel free to use it if it could be of a little help! It should produce almost identical output as librosa according to my simple tests. I just collected things from internet too :)</p>\n<p>BTW, for some reason, the last <code>power_to_db</code> step can also be faster.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F0c5522e414a0bdc294c663fee9563f98%2FScreenshot%20from%202021-01-09%2015-12-13.png?generation=1610226880419856&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146608,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/09/2021 21:59:39",
          "content": "<p>Wow, that's great. Understanding/knowing the background of this computations definitely makes everything much more smooth. </p>\n<p>Also, I'm trying to adapt your code to my pipeline. In terms of augmentation, let's say I want to inject some noise to signal. </p>\n<p>As I was searching about this topic, I came to an idea that we can add noise in two ways:</p>\n<ol>\n<li>Adding noise to raw signal like just after reading the audio file.</li>\n<li>Adding noise to the features like melspectrogram (I did not convince myself about this one yet)</li>\n</ol>\n<p>Where should I apply this noise? I will use your variable names here.</p>\n<p>A. wav<br>\nB. fft_windows<br>\nC. s2</p>\n<p>Thanks a lot again 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146691,
          "author_name": "barnwellguy",
          "author_url": "",
          "post_date": "01/10/2021 00:25:48",
          "content": "<p>I didn't augment the data nearly at all so far in this competition, so I don't have much practical advice. </p>\n<p>But, if we want to go approach 1, which perhaps is most natural, then, of course, injecting noise to wave and do all steps of computation will work. If that is too slow, and we want exactly the same result as going through all steps, then, we can select a collection of typical noises that contains not any target bird call in it and prercompute STFT/fft_windows for it. When we want to inject any of this noise into any sample, we just add the STFT of the noise to the STFT of the sample. </p>\n<p>Since STFT is linear, STFT(sample + noise) = STFT(sample) + STFT(noise)</p>\n<p>So, after precompuation of STFT(sample) and STFT(noise), we don't need to do the most expensive STFT for every sample with noise when we generate data on the go.</p>\n<p>About approach 2, adding noise to <code>s2</code>, or mix sample's <code>s2</code> with other sample's <code>s2</code> is like those typical \"mix-up\" method image processing project does, right? Although, the generated images may not have any corresponding realistic sound which is perhaps not ideal, but it still helps the model to avoid learning trivial spurious characteristics, right? e.g., it is species 5 if the 3rd pixel in the bottom line is between 5.5 and 6.5. Especially, we only have 50 positive data for each species.</p>\n<p>Above are just my 2 cents, without much consideration of this specific context 🤝</p>\n<p>We can discuss more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147980,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/10/2021 20:37:32",
          "content": "<p>Hi again, </p>\n<p>First of all, I just want to say thank you. In this phase of the competition, I finally managed to pass 87.6 (the highest scoring public notebook) with your support.<br>\nI went with the option 1 and applied noise to raw signal.  Everything so far seems to work well in my current setup. Things started to make much more sense.</p>\n<p>Thanks a lot again 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148205,
          "author_name": "barnwellguy",
          "author_url": "",
          "post_date": "01/11/2021 02:01:30",
          "content": "<p>Congratulations! Happy to hear it helped!<br>\nIt's not easy to beat ensemble submissions! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149337,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/11/2021 19:15:10",
          "content": "<p>Thanks a lot 😊 </p>\n<p>My solution is also based on ensembles, but there are some improvements on the base models thanks to your ideas.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1141385": "Hi everyone,\n\nI was precomputing the audio files and save them to the file. Then, I was loading them in the dataloader. But I have added a data augmentation to my pipeline and because of that I can not do precompute & load method since I have to apply data augmentation in the process itself. Before the data augmentation my single epoch was taking a minute or so. But now with data augmentation it takes 6 to 7 minutes. \n\nFirst of all, is my approach wrong? \nSecondly, what can I do to speed up this process?\n\nThanks.",
    "1142703": "Hi, @snnclsr , good question! I was also thinking about this problem. Did some experiments, and hopefully a bit helpful.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2Fc01dd3b559520a7c99a061806df8a673%2FScreenshot%20from%202021-01-07%2008-39-42.png?generation=1610030415862037&alt=media)\n\nThe above notebook demonstrates couple of things,\n- The melspectrogram computation is mainly three parts: STFT, absolute&power, convert to Mel scale, and optionally covnert to db (which is taking a log).\n- the most time consuming STFT is a linear operation, in the sense that, if we want to add noise to the original input, it's equivalent to add the STFT of noise to the STFT of the original input. So, if noise are from a dictionary of audio clips, then, we can precompute STFT of original inputs and noises, the STFT of their random combination does not needs to be recomputed.\n- The absolute&power step is also slow using the original implementation. But if we keep the power to be 2, then, the provided numba implementation is much faster than the numpy one, with minimum discrepancy. I learnt this from [stackoverflow](https://stackoverflow.com/questions/30437947/most-memory-efficient-way-to-compute-abs2-of-complex-numpy-ndarray) and there are couple of more implementations.",
    "1146510": "Hi @barnwellguy , Thanks a lot for your detailed and great answer! Your answer lighten some dark places in my mind. So instead of using the librosa's melspectrogram function, can we use your version ? I will definitely try btw. \n\nAlso I don't have any prior experience in audio processing. So everything is just new to me :)\nThanks a lot again.",
    "1146556": "I am glad that you liked it. Feel free to use it if it could be of a little help! It should produce almost identical output as librosa according to my simple tests. I just collected things from internet too :)\n\nBTW, for some reason, the last `power_to_db` step can also be faster.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F0c5522e414a0bdc294c663fee9563f98%2FScreenshot%20from%202021-01-09%2015-12-13.png?generation=1610226880419856&alt=media)",
    "1146608": "Wow, that's great. Understanding/knowing the background of this computations definitely makes everything much more smooth. \n\nAlso, I'm trying to adapt your code to my pipeline. In terms of augmentation, let's say I want to inject some noise to signal. \n\nAs I was searching about this topic, I came to an idea that we can add noise in two ways:\n\n1. Adding noise to raw signal like just after reading the audio file.\n2. Adding noise to the features like melspectrogram (I did not convince myself about this one yet)\n\nWhere should I apply this noise? I will use your variable names here.\n\nA. wav\nB. fft_windows\nC. s2\n\nThanks a lot again 🙏",
    "1146691": "I didn't augment the data nearly at all so far in this competition, so I don't have much practical advice. \n\nBut, if we want to go approach 1, which perhaps is most natural, then, of course, injecting noise to wave and do all steps of computation will work. If that is too slow, and we want exactly the same result as going through all steps, then, we can select a collection of typical noises that contains not any target bird call in it and prercompute STFT/fft_windows for it. When we want to inject any of this noise into any sample, we just add the STFT of the noise to the STFT of the sample. \n\nSince STFT is linear, STFT(sample + noise) = STFT(sample) + STFT(noise)\n\nSo, after precompuation of STFT(sample) and STFT(noise), we don't need to do the most expensive STFT for every sample with noise when we generate data on the go.\n\nAbout approach 2, adding noise to `s2`, or mix sample's `s2` with other sample's `s2` is like those typical \"mix-up\" method image processing project does, right? Although, the generated images may not have any corresponding realistic sound which is perhaps not ideal, but it still helps the model to avoid learning trivial spurious characteristics, right? e.g., it is species 5 if the 3rd pixel in the bottom line is between 5.5 and 6.5. Especially, we only have 50 positive data for each species.\n\nAbove are just my 2 cents, without much consideration of this specific context 🤝\n\nWe can discuss more.",
    "1147980": "Hi again, \n\nFirst of all, I just want to say thank you. In this phase of the competition, I finally managed to pass 87.6 (the highest scoring public notebook) with your support.\nI went with the option 1 and applied noise to raw signal.  Everything so far seems to work well in my current setup. Things started to make much more sense.\n\nThanks a lot again 🙏",
    "1148205": "Congratulations! Happy to hear it helped!\nIt's not easy to beat ensemble submissions!",
    "1149337": "Thanks a lot 😊 \n\nMy solution is also based on ensembles, but there are some improvements on the base models thanks to your ideas."
  },
  "source": "meta"
}