{
  "id": 234851,
  "title": "Augmentations BirdCLEF 2021",
  "url": "/competitions/birdclef-2021/discussion/234851",
  "author_name": "",
  "post_date": "2021-04-26T14:15:32.336290500Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi folks!<br>\nWhich strategy did you use to add some augmentations to this competition?<br>\nI'm using this wonderful <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">augs</a> from last birds competition. <br>\nThat's works, but i had some problem - one epoch training increased from ~10 mins to ~1 hour. Maybe that's some mistakes in my code and its some ways to optimize this. <br>\nSo maybe someone using albumentations or tf augmentations? <br>\nThx advice. </p>",
  "messages": [
    {
      "id": "1285041",
      "postDate": "04/26/2021 14:15:32",
      "content": "<p>Hi folks!<br>\nWhich strategy did you use to add some augmentations to this competition?<br>\nI'm using this wonderful <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">augs</a> from last birds competition. <br>\nThat's works, but i had some problem - one epoch training increased from ~10 mins to ~1 hour. Maybe that's some mistakes in my code and its some ways to optimize this. <br>\nSo maybe someone using albumentations or tf augmentations? <br>\nThx advice. </p>",
      "rawMarkdown": "Hi folks!\nWhich strategy did you use to add some augmentations to this competition?\nI'm using this wonderful [augs](https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english) from last birds competition. \nThat's works, but i had some problem - one epoch training increased from ~10 mins to ~1 hour. Maybe that's some mistakes in my code and its some ways to optimize this. \nSo maybe someone using albumentations or tf augmentations? \nThx advice.",
      "votes": null
    },
    {
      "id": "1285162",
      "postDate": "04/26/2021 16:35:21",
      "content": "<p>I recently joined this competition and have very less knowledge. But by reading through previous competitions I got to know that by converting the audio files to images -&gt; storing them -&gt; reloading them makes things faster rather than preparing the images on fly..<br>\nI am not sure if the augs can be applied directly on images or not. </p>",
      "rawMarkdown": "I recently joined this competition and have very less knowledge. But by reading through previous competitions I got to know that by converting the audio files to images -> storing them -> reloading them makes things faster rather than preparing the images on fly..\nI am not sure if the augs can be applied directly on images or not.",
      "votes": null
    },
    {
      "id": "1286469",
      "postDate": "04/28/2021 03:30:50",
      "content": "<p>Try <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">torch-audiomentations</a>, which could run audio augmentation on GPU.</p>",
      "rawMarkdown": "Try [torch-audiomentations](https://github.com/asteroid-team/torch-audiomentations), which could run audio augmentation on GPU.",
      "votes": null
    },
    {
      "id": "1304896",
      "postDate": "05/13/2021 01:07:16",
      "content": "<p>Hello, I would like to ask you which image augmentation is available in this task.looking forward to your reply.</p>",
      "rawMarkdown": "Hello, I would like to ask you which image augmentation is available in this task.looking forward to your reply.",
      "votes": null
    },
    {
      "id": "1305133",
      "postDate": "05/13/2021 05:36:48",
      "content": "<p>To avoid augmentations slowing your training cycle - you should use multithreaded data loading, for example - <a href=\"https://pytorch.org/docs/stable/data.html#multi-process-data-loading\" target=\"_blank\">https://pytorch.org/docs/stable/data.html#multi-process-data-loading</a></p>\n<p>Then as long as you have enough CPU cores to prepare next batch of data while current one is being processed on GPU - your slowdown would be zero.</p>",
      "rawMarkdown": "To avoid augmentations slowing your training cycle - you should use multithreaded data loading, for example - https://pytorch.org/docs/stable/data.html#multi-process-data-loading\n\nThen as long as you have enough CPU cores to prepare next batch of data while current one is being processed on GPU - your slowdown would be zero.",
      "votes": null
    },
    {
      "id": "1305284",
      "postDate": "05/13/2021 07:36:25",
      "content": "<p>Please check audiomentations, torch-audiomentations</p>",
      "rawMarkdown": "Please check audiomentations, torch-audiomentations",
      "votes": null
    },
    {
      "id": "1318770",
      "postDate": "05/22/2021 15:04:02",
      "content": "<p>Hey folks! I wanted to do some (background) noise augmentation to the train_short_audio files and to keep it simpel, I'm using the train_soundscapes as the noise. First, based on the train_soundscape_labels, I removed all the segments from the train_soundscapes files, where a bird is singing. This left me with (somewhat) pure train_soundscapes noise-only files. </p>\n<p>Now, inspired by the comments, I'm using audiomentations' AddBackgroundNoise, to add the noise from the train_soundscapes to the short_train_audio. I've set p=1, so it's always added (as I'm keeping the originals). </p>\n<p>However, I don't know how to set the min_snr_in_db and max_snr_in_db meaningfully (based on anything theoretical). I can observe that, the lower the max_snr_in_db is, the higher the resulting noise. I guess, without these parameters set, they will be set based on how loud the signal already is from the beginning? (I.e. if the bird is very low to begin with, it might not make sense to just add very loud noise to it, as it will drown completely). But I'm not sure of this.  <br>\nCan someone explain the effect of setting the min_snr_in_db both high and low and the effect of setting the max_snr_in_db both high and low? Thanks! :-)</p>\n<p>EDIT: Reading the source code of audiomentations, the min and max values are simply to set a range, from within a uniform sample will be drawn to be the signal to noise for that particular augmentation. </p>\n<p><code>self.parameters[\"snr_in_db\"] = random.uniform(\n  self.min_snr_in_db, self.max_snr_in_db</code></p>\n<p>But I'm still curious, how we should choose the signal to noise ratio. </p>",
      "rawMarkdown": "Hey folks! I wanted to do some (background) noise augmentation to the train_short_audio files and to keep it simpel, I'm using the train_soundscapes as the noise. First, based on the train_soundscape_labels, I removed all the segments from the train_soundscapes files, where a bird is singing. This left me with (somewhat) pure train_soundscapes noise-only files. \n\nNow, inspired by the comments, I'm using audiomentations' AddBackgroundNoise, to add the noise from the train_soundscapes to the short_train_audio. I've set p=1, so it's always added (as I'm keeping the originals). \n\nHowever, I don't know how to set the min_snr_in_db and max_snr_in_db meaningfully (based on anything theoretical). I can observe that, the lower the max_snr_in_db is, the higher the resulting noise. I guess, without these parameters set, they will be set based on how loud the signal already is from the beginning? (I.e. if the bird is very low to begin with, it might not make sense to just add very loud noise to it, as it will drown completely). But I'm not sure of this.  \nCan someone explain the effect of setting the min_snr_in_db both high and low and the effect of setting the max_snr_in_db both high and low? Thanks! :-)\n\nEDIT: Reading the source code of audiomentations, the min and max values are simply to set a range, from within a uniform sample will be drawn to be the signal to noise for that particular augmentation. \n\n`self.parameters[\"snr_in_db\"] = random.uniform(\n  self.min_snr_in_db, self.max_snr_in_db`\n\nBut I'm still curious, how we should choose the signal to noise ratio.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1285162,
      "author_name": "nitindatta",
      "author_url": "",
      "post_date": "04/26/2021 16:35:21",
      "content": "<p>I recently joined this competition and have very less knowledge. But by reading through previous competitions I got to know that by converting the audio files to images -&gt; storing them -&gt; reloading them makes things faster rather than preparing the images on fly..<br>\nI am not sure if the augs can be applied directly on images or not. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1304896,
          "author_name": "majunfu",
          "author_url": "",
          "post_date": "05/13/2021 01:07:16",
          "content": "<p>Hello, I would like to ask you which image augmentation is available in this task.looking forward to your reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305284,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "05/13/2021 07:36:25",
          "content": "<p>Please check audiomentations, torch-audiomentations</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1286469,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "04/28/2021 03:30:50",
      "content": "<p>Try <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\">torch-audiomentations</a>, which could run audio augmentation on GPU.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1305133,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "05/13/2021 05:36:48",
      "content": "<p>To avoid augmentations slowing your training cycle - you should use multithreaded data loading, for example - <a href=\"https://pytorch.org/docs/stable/data.html#multi-process-data-loading\" target=\"_blank\">https://pytorch.org/docs/stable/data.html#multi-process-data-loading</a></p>\n<p>Then as long as you have enough CPU cores to prepare next batch of data while current one is being processed on GPU - your slowdown would be zero.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1318770,
      "author_name": "tobiasbonnesen",
      "author_url": "",
      "post_date": "05/22/2021 15:04:02",
      "content": "<p>Hey folks! I wanted to do some (background) noise augmentation to the train_short_audio files and to keep it simpel, I'm using the train_soundscapes as the noise. First, based on the train_soundscape_labels, I removed all the segments from the train_soundscapes files, where a bird is singing. This left me with (somewhat) pure train_soundscapes noise-only files. </p>\n<p>Now, inspired by the comments, I'm using audiomentations' AddBackgroundNoise, to add the noise from the train_soundscapes to the short_train_audio. I've set p=1, so it's always added (as I'm keeping the originals). </p>\n<p>However, I don't know how to set the min_snr_in_db and max_snr_in_db meaningfully (based on anything theoretical). I can observe that, the lower the max_snr_in_db is, the higher the resulting noise. I guess, without these parameters set, they will be set based on how loud the signal already is from the beginning? (I.e. if the bird is very low to begin with, it might not make sense to just add very loud noise to it, as it will drown completely). But I'm not sure of this.  <br>\nCan someone explain the effect of setting the min_snr_in_db both high and low and the effect of setting the max_snr_in_db both high and low? Thanks! :-)</p>\n<p>EDIT: Reading the source code of audiomentations, the min and max values are simply to set a range, from within a uniform sample will be drawn to be the signal to noise for that particular augmentation. </p>\n<p><code>self.parameters[\"snr_in_db\"] = random.uniform(\n  self.min_snr_in_db, self.max_snr_in_db</code></p>\n<p>But I'm still curious, how we should choose the signal to noise ratio. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1285041": "Hi folks!\nWhich strategy did you use to add some augmentations to this competition?\nI'm using this wonderful [augs](https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english) from last birds competition. \nThat's works, but i had some problem - one epoch training increased from ~10 mins to ~1 hour. Maybe that's some mistakes in my code and its some ways to optimize this. \nSo maybe someone using albumentations or tf augmentations? \nThx advice.",
    "1285162": "I recently joined this competition and have very less knowledge. But by reading through previous competitions I got to know that by converting the audio files to images -> storing them -> reloading them makes things faster rather than preparing the images on fly..\nI am not sure if the augs can be applied directly on images or not.",
    "1286469": "Try [torch-audiomentations](https://github.com/asteroid-team/torch-audiomentations), which could run audio augmentation on GPU.",
    "1304896": "Hello, I would like to ask you which image augmentation is available in this task.looking forward to your reply.",
    "1305133": "To avoid augmentations slowing your training cycle - you should use multithreaded data loading, for example - https://pytorch.org/docs/stable/data.html#multi-process-data-loading\n\nThen as long as you have enough CPU cores to prepare next batch of data while current one is being processed on GPU - your slowdown would be zero.",
    "1305284": "Please check audiomentations, torch-audiomentations",
    "1318770": "Hey folks! I wanted to do some (background) noise augmentation to the train_short_audio files and to keep it simpel, I'm using the train_soundscapes as the noise. First, based on the train_soundscape_labels, I removed all the segments from the train_soundscapes files, where a bird is singing. This left me with (somewhat) pure train_soundscapes noise-only files. \n\nNow, inspired by the comments, I'm using audiomentations' AddBackgroundNoise, to add the noise from the train_soundscapes to the short_train_audio. I've set p=1, so it's always added (as I'm keeping the originals). \n\nHowever, I don't know how to set the min_snr_in_db and max_snr_in_db meaningfully (based on anything theoretical). I can observe that, the lower the max_snr_in_db is, the higher the resulting noise. I guess, without these parameters set, they will be set based on how loud the signal already is from the beginning? (I.e. if the bird is very low to begin with, it might not make sense to just add very loud noise to it, as it will drown completely). But I'm not sure of this.  \nCan someone explain the effect of setting the min_snr_in_db both high and low and the effect of setting the max_snr_in_db both high and low? Thanks! :-)\n\nEDIT: Reading the source code of audiomentations, the min and max values are simply to set a range, from within a uniform sample will be drawn to be the signal to noise for that particular augmentation. \n\n`self.parameters[\"snr_in_db\"] = random.uniform(\n  self.min_snr_in_db, self.max_snr_in_db`\n\nBut I'm still curious, how we should choose the signal to noise ratio."
  },
  "source": "meta"
}