{
  "id": 178917,
  "title": "What are the Augmentations that have worked for you?",
  "url": "/competitions/birdsong-recognition/discussion/178917",
  "author_name": "",
  "post_date": "2020-08-31T21:14:29.903848900Z",
  "votes": 16,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Finally after making my baseline work , I am moving to improve it , I have tried a lot of augs now and for me , augmenting audio rather than spectogram works better . I am trying to find a way to use cutmix and mixup directly on audio as well .</p>\n<p>It would be great if the community can share what augmentations have worked for them and what have not</p>\n<p>Here is the list of augs that have worked for me :-</p>\n<ul>\n<li>Adding Noise</li>\n<li>Cut out</li>\n<li>Adding noise from random data like freeaudiotagging data</li>\n<li>Time Stretch</li>\n</ul>\n<p>The frequency masking and time masking doesn't seem to be working for me.</p>\n<p>Also I feel doing TTA with audio transforms will be a better idea than with image transforms , has anyone tried TTA with audio transforms??</p>",
  "messages": [
    {
      "id": "993344",
      "postDate": "08/31/2020 21:14:29",
      "content": "<p>Finally after making my baseline work , I am moving to improve it , I have tried a lot of augs now and for me , augmenting audio rather than spectogram works better . I am trying to find a way to use cutmix and mixup directly on audio as well .</p>\n<p>It would be great if the community can share what augmentations have worked for them and what have not</p>\n<p>Here is the list of augs that have worked for me :-</p>\n<ul>\n<li>Adding Noise</li>\n<li>Cut out</li>\n<li>Adding noise from random data like freeaudiotagging data</li>\n<li>Time Stretch</li>\n</ul>\n<p>The frequency masking and time masking doesn't seem to be working for me.</p>\n<p>Also I feel doing TTA with audio transforms will be a better idea than with image transforms , has anyone tried TTA with audio transforms??</p>",
      "rawMarkdown": "Finally after making my baseline work , I am moving to improve it , I have tried a lot of augs now and for me , augmenting audio rather than spectogram works better . I am trying to find a way to use cutmix and mixup directly on audio as well .\n\nIt would be great if the community can share what augmentations have worked for them and what have not\n\nHere is the list of augs that have worked for me :-\n* Adding Noise\n* Cut out\n* Adding noise from random data like freeaudiotagging data\n* Time Stretch\n\nThe frequency masking and time masking doesn't seem to be working for me.\n\nAlso I feel doing TTA with audio transforms will be a better idea than with image transforms , has anyone tried TTA with audio transforms??",
      "votes": null
    },
    {
      "id": "993370",
      "postDate": "08/31/2020 21:54:04",
      "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> Can you please share what noise level do you add the Freesound Data? I found that when I added it, my score got worse. Do you just do y = y + noise, or must be doing y = y + noise*0.1  …</p>\n<p>Also, can you comment on time stretch? I found using librosa implementation is very slow. What are you using?</p>\n<p>I can concur with you that adding noise seems helpful</p>",
      "rawMarkdown": "tanulsingh077 Can you please share what noise level do you add the Freesound Data? I found that when I added it, my score got worse. Do you just do y = y + noise, or must be doing y = y + noise*0.1  ...\n\nAlso, can you comment on time stretch? I found using librosa implementation is very slow. What are you using?\n\nI can concur with you that adding noise seems helpful",
      "votes": null
    },
    {
      "id": "993887",
      "postDate": "09/01/2020 07:50:38",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> I am doing y = y + noise*0.1…..  also I have seen that adding background noise works better than adding gaussian noise , what I do is I add files from free audio tagging comp to the audio files here. I am doing all the augmentations using the same techniques I have shared in my kernel here : <a href=\"https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio\" target=\"_blank\">https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio</a></p>",
      "rawMarkdown": "Hello @returnofsputnik I am doing y = y + noise*0.1.....  also I have seen that adding background noise works better than adding gaussian noise , what I do is I add files from free audio tagging comp to the audio files here. I am doing all the augmentations using the same techniques I have shared in my kernel here : https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio",
      "votes": null
    },
    {
      "id": "993893",
      "postDate": "09/01/2020 07:59:39",
      "content": "<p>Is not that slow? I used all of the augmentations from your kernel but that made things slow while training. The issue is that when we are using Librosa it takes time to do the operation so instead of that it is a better idea to do everything using numpy or some vanilla approach. For example, adding audio from the free-sound competition can be done in two steps, </p>\n<ol>\n<li>we read the audio files and store them as a numpy array using Librosa separately. </li>\n<li>Then use the array of sound to get any random audio and add that into the original bird audio. </li>\n</ol>\n<p>This makes things fast. </p>",
      "rawMarkdown": "Is not that slow? I used all of the augmentations from your kernel but that made things slow while training. The issue is that when we are using Librosa it takes time to do the operation so instead of that it is a better idea to do everything using numpy or some vanilla approach. For example, adding audio from the free-sound competition can be done in two steps, \n\n1. we read the audio files and store them as a numpy array using Librosa separately. \n2. Then use the array of sound to get any random audio and add that into the original bird audio. \n\nThis makes things fast.",
      "votes": null
    },
    {
      "id": "994046",
      "postDate": "09/01/2020 10:28:11",
      "content": "<p><a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> small tip:</p>\n<blockquote>\n  <p>first pick the 5 seconds audio, then apply transform, it will be much faster this way.</p>\n</blockquote>",
      "rawMarkdown": "urvishp80 small tip:\n> first pick the 5 seconds audio, then apply transform, it will be much faster this way.",
      "votes": null
    },
    {
      "id": "994055",
      "postDate": "09/01/2020 10:35:32",
      "content": "<p>will try this way for sure. Thanks for the tip. <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> </p>",
      "rawMarkdown": "will try this way for sure. Thanks for the tip. @rohitsingh9990",
      "votes": null
    },
    {
      "id": "994684",
      "postDate": "09/01/2020 20:02:15",
      "content": "<p><a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> Rohit already answered your query haha</p>",
      "rawMarkdown": "urvishp80 Rohit already answered your query haha",
      "votes": null
    },
    {
      "id": "994894",
      "postDate": "09/02/2020 03:40:32",
      "content": "<p>I think mixup in audio looks something like:</p>\n<pre><code>mixup_audio = 0.5*audio_1 + 0.5*audio_2\n</code></pre>\n<p>Basically audio overlapping</p>",
      "rawMarkdown": "I think mixup in audio looks something like:\n```\nmixup_audio = 0.5*audio_1 + 0.5*audio_2\n```\nBasically audio overlapping",
      "votes": null
    },
    {
      "id": "995754",
      "postDate": "09/02/2020 18:36:56",
      "content": "<p>As Cut out do you mean with that the known image augmentation?</p>",
      "rawMarkdown": "As Cut out do you mean with that the known image augmentation?",
      "votes": null
    },
    {
      "id": "996216",
      "postDate": "09/03/2020 06:38:14",
      "content": "<p>Mixup works really well for me. I also can see many audio classification paper use it.</p>",
      "rawMarkdown": "Mixup works really well for me. I also can see many audio classification paper use it.",
      "votes": null
    },
    {
      "id": "996563",
      "postDate": "09/03/2020 11:51:57",
      "content": "<p>Are you mixup melspec or audio? I cant get it to work, maybe i implement it wrong</p>",
      "rawMarkdown": "Are you mixup melspec or audio? I cant get it to work, maybe i implement it wrong",
      "votes": null
    },
    {
      "id": "996786",
      "postDate": "09/03/2020 14:47:53",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/quandapro\" target=\"_blank\">@quandapro</a>  yes that's right  but the tricky part is how do we integrate with other augs and use simultaneously because mixup also augments the labels as well and use multiple audios from same batch . Thus even if you implement mixup the function will have to return labels as well as mixedaudio while other augs only return audio file</p>",
      "rawMarkdown": "Hey @quandapro  yes that's right  but the tricky part is how do we integrate with other augs and use simultaneously because mixup also augments the labels as well and use multiple audios from same batch . Thus even if you implement mixup the function will have to return labels as well as mixedaudio while other augs only return audio file",
      "votes": null
    },
    {
      "id": "996787",
      "postDate": "09/03/2020 14:48:16",
      "content": "<p>Yes I just randomly make a segment of audio zero</p>",
      "rawMarkdown": "Yes I just randomly make a segment of audio zero",
      "votes": null
    },
    {
      "id": "998340",
      "postDate": "09/04/2020 17:08:29",
      "content": "<p>same question here! My model cant converge well if adding mixup.</p>",
      "rawMarkdown": "same question here! My model cant converge well if adding mixup.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 993370,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "08/31/2020 21:54:04",
      "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> Can you please share what noise level do you add the Freesound Data? I found that when I added it, my score got worse. Do you just do y = y + noise, or must be doing y = y + noise*0.1  …</p>\n<p>Also, can you comment on time stretch? I found using librosa implementation is very slow. What are you using?</p>\n<p>I can concur with you that adding noise seems helpful</p>",
      "votes": null,
      "replies": [
        {
          "id": 993887,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "09/01/2020 07:50:38",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> I am doing y = y + noise*0.1…..  also I have seen that adding background noise works better than adding gaussian noise , what I do is I add files from free audio tagging comp to the audio files here. I am doing all the augmentations using the same techniques I have shared in my kernel here : <a href=\"https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio\" target=\"_blank\">https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 993893,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "09/01/2020 07:59:39",
          "content": "<p>Is not that slow? I used all of the augmentations from your kernel but that made things slow while training. The issue is that when we are using Librosa it takes time to do the operation so instead of that it is a better idea to do everything using numpy or some vanilla approach. For example, adding audio from the free-sound competition can be done in two steps, </p>\n<ol>\n<li>we read the audio files and store them as a numpy array using Librosa separately. </li>\n<li>Then use the array of sound to get any random audio and add that into the original bird audio. </li>\n</ol>\n<p>This makes things fast. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 994046,
          "author_name": "rohitsingh9990",
          "author_url": "",
          "post_date": "09/01/2020 10:28:11",
          "content": "<p><a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> small tip:</p>\n<blockquote>\n  <p>first pick the 5 seconds audio, then apply transform, it will be much faster this way.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 994055,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "09/01/2020 10:35:32",
          "content": "<p>will try this way for sure. Thanks for the tip. <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 994684,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "09/01/2020 20:02:15",
          "content": "<p><a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> Rohit already answered your query haha</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 994894,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "09/02/2020 03:40:32",
      "content": "<p>I think mixup in audio looks something like:</p>\n<pre><code>mixup_audio = 0.5*audio_1 + 0.5*audio_2\n</code></pre>\n<p>Basically audio overlapping</p>",
      "votes": null,
      "replies": [
        {
          "id": 996786,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "09/03/2020 14:47:53",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/quandapro\" target=\"_blank\">@quandapro</a>  yes that's right  but the tricky part is how do we integrate with other augs and use simultaneously because mixup also augments the labels as well and use multiple audios from same batch . Thus even if you implement mixup the function will have to return labels as well as mixedaudio while other augs only return audio file</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 995754,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "09/02/2020 18:36:56",
      "content": "<p>As Cut out do you mean with that the known image augmentation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 996787,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "09/03/2020 14:48:16",
          "content": "<p>Yes I just randomly make a segment of audio zero</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 996216,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "09/03/2020 06:38:14",
      "content": "<p>Mixup works really well for me. I also can see many audio classification paper use it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 996563,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "09/03/2020 11:51:57",
          "content": "<p>Are you mixup melspec or audio? I cant get it to work, maybe i implement it wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998340,
          "author_name": "fiyeroleung",
          "author_url": "",
          "post_date": "09/04/2020 17:08:29",
          "content": "<p>same question here! My model cant converge well if adding mixup.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "993344": "Finally after making my baseline work , I am moving to improve it , I have tried a lot of augs now and for me , augmenting audio rather than spectogram works better . I am trying to find a way to use cutmix and mixup directly on audio as well .\n\nIt would be great if the community can share what augmentations have worked for them and what have not\n\nHere is the list of augs that have worked for me :-\n* Adding Noise\n* Cut out\n* Adding noise from random data like freeaudiotagging data\n* Time Stretch\n\nThe frequency masking and time masking doesn't seem to be working for me.\n\nAlso I feel doing TTA with audio transforms will be a better idea than with image transforms , has anyone tried TTA with audio transforms??",
    "993370": "tanulsingh077 Can you please share what noise level do you add the Freesound Data? I found that when I added it, my score got worse. Do you just do y = y + noise, or must be doing y = y + noise*0.1  ...\n\nAlso, can you comment on time stretch? I found using librosa implementation is very slow. What are you using?\n\nI can concur with you that adding noise seems helpful",
    "993887": "Hello @returnofsputnik I am doing y = y + noise*0.1.....  also I have seen that adding background noise works better than adding gaussian noise , what I do is I add files from free audio tagging comp to the audio files here. I am doing all the augmentations using the same techniques I have shared in my kernel here : https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio",
    "993893": "Is not that slow? I used all of the augmentations from your kernel but that made things slow while training. The issue is that when we are using Librosa it takes time to do the operation so instead of that it is a better idea to do everything using numpy or some vanilla approach. For example, adding audio from the free-sound competition can be done in two steps, \n\n1. we read the audio files and store them as a numpy array using Librosa separately. \n2. Then use the array of sound to get any random audio and add that into the original bird audio. \n\nThis makes things fast.",
    "994046": "urvishp80 small tip:\n> first pick the 5 seconds audio, then apply transform, it will be much faster this way.",
    "994055": "will try this way for sure. Thanks for the tip. @rohitsingh9990",
    "994684": "urvishp80 Rohit already answered your query haha",
    "994894": "I think mixup in audio looks something like:\n```\nmixup_audio = 0.5*audio_1 + 0.5*audio_2\n```\nBasically audio overlapping",
    "995754": "As Cut out do you mean with that the known image augmentation?",
    "996216": "Mixup works really well for me. I also can see many audio classification paper use it.",
    "996563": "Are you mixup melspec or audio? I cant get it to work, maybe i implement it wrong",
    "996786": "Hey @quandapro  yes that's right  but the tricky part is how do we integrate with other augs and use simultaneously because mixup also augments the labels as well and use multiple audios from same batch . Thus even if you implement mixup the function will have to return labels as well as mixedaudio while other augs only return audio file",
    "996787": "Yes I just randomly make a segment of audio zero",
    "998340": "same question here! My model cant converge well if adding mixup."
  },
  "source": "meta"
}