{
  "id": 444979,
  "title": "Are Data Augmentations Helpful?",
  "url": "/competitions/bengaliai-speech/discussion/444979",
  "author_name": "",
  "post_date": "2023-10-04T14:57:27.884222400Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Has anyone had any luck training with data augmentation strategies? I know data augmentation is really helpful in CV (unless some new research paper I haven't read has overturned that), but I was surprised (maybe due to my naivety in DL audio) to find out audio had so many augmentation types.</p>\n<p>I slapped together a notebook, basically replicating <a href=\"https://www.youtube.com/playlist?list=PL-wATfeyAMNoR4aqS-Fv0GRmS6bx5RtTW\" target=\"_blank\">The Sound of AI's</a> tutorial. For a sample on LibriSpeech, seems the distortions aren't too bad, but there are enough artifacts for me to question efficacy. CTC loss seems to decrease when using them, but I haven't trained on a mass scale.</p>\n<p>Maybe someone who's a native speaker can attest to the quality of the Bengali audio when transformed?</p>\n<p><a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/msthil/bengali-data-augmentation\" target=\"_blank\">https://www.kaggle.com/code/msthil/bengali-data-augmentation</a> </p>",
  "messages": [
    {
      "id": "2467447",
      "postDate": "10/04/2023 14:57:27",
      "content": "<p>Has anyone had any luck training with data augmentation strategies? I know data augmentation is really helpful in CV (unless some new research paper I haven't read has overturned that), but I was surprised (maybe due to my naivety in DL audio) to find out audio had so many augmentation types.</p>\n<p>I slapped together a notebook, basically replicating <a href=\"https://www.youtube.com/playlist?list=PL-wATfeyAMNoR4aqS-Fv0GRmS6bx5RtTW\" target=\"_blank\">The Sound of AI's</a> tutorial. For a sample on LibriSpeech, seems the distortions aren't too bad, but there are enough artifacts for me to question efficacy. CTC loss seems to decrease when using them, but I haven't trained on a mass scale.</p>\n<p>Maybe someone who's a native speaker can attest to the quality of the Bengali audio when transformed?</p>\n<p><a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/msthil/bengali-data-augmentation\" target=\"_blank\">https://www.kaggle.com/code/msthil/bengali-data-augmentation</a> </p>",
      "rawMarkdown": "Has anyone had any luck training with data augmentation strategies? I know data augmentation is really helpful in CV (unless some new research paper I haven't read has overturned that), but I was surprised (maybe due to my naivety in DL audio) to find out audio had so many augmentation types.\n\nI slapped together a notebook, basically replicating [The Sound of AI's](https://www.youtube.com/playlist?list=PL-wATfeyAMNoR4aqS-Fv0GRmS6bx5RtTW) tutorial. For a sample on LibriSpeech, seems the distortions aren't too bad, but there are enough artifacts for me to question efficacy. CTC loss seems to decrease when using them, but I haven't trained on a mass scale.\n\nMaybe someone who's a native speaker can attest to the quality of the Bengali audio when transformed?\n\n[https://www.kaggle.com/code/msthil/bengali-data-augmentation ](url)",
      "votes": null
    },
    {
      "id": "2467510",
      "postDate": "10/04/2023 16:05:06",
      "content": "<p>Data augmentation can indeed be beneficial for audio data, introducing variations to improve model robustness. However, it's crucial to strike a balance, as excessive augmentation may introduce artifacts. If you've had success or insights, please consider upvoting to share your experiences and foster discussions on effective audio data augmentation strategies. Thank you!</p>",
      "rawMarkdown": "Data augmentation can indeed be beneficial for audio data, introducing variations to improve model robustness. However, it's crucial to strike a balance, as excessive augmentation may introduce artifacts. If you've had success or insights, please consider upvoting to share your experiences and foster discussions on effective audio data augmentation strategies. Thank you!",
      "votes": null
    },
    {
      "id": "2473015",
      "postDate": "10/07/2023 19:31:06",
      "content": "<p>May I ask, did you get any gains with this approach?</p>",
      "rawMarkdown": "May I ask, did you get any gains with this approach?",
      "votes": null
    },
    {
      "id": "2473282",
      "postDate": "10/08/2023 06:34:57",
      "content": "<p>Unfortunately, I haven't had the time to thoroughly try this, however, I did do a quick and dirt training of about 1000 steps using the below augmentation parameters and the loss decreased out from (what I assume was) a local minimum I seemed to have been caught in. That decrease though wasn't too crazy. Hoping to have more time over the next few days to go crazy on a bigger GPU. Please do share though if you manage to see some gains.</p>\n<pre><code>augment = Compose([\n            AddGaussianNoise(min_amplitude=, max_amplitude=, p=),\n            TimeStretch(min_rate=, max_rate=, p=),\n            PitchShift(min_semitones=-, max_semitones=, p=),\n            Shift(p=),\n            Trim(top_db=, p=)\n        ])\n</code></pre>",
      "rawMarkdown": "Unfortunately, I haven't had the time to thoroughly try this, however, I did do a quick and dirt training of about 1000 steps using the below augmentation parameters and the loss decreased out from (what I assume was) a local minimum I seemed to have been caught in. That decrease though wasn't too crazy. Hoping to have more time over the next few days to go crazy on a bigger GPU. Please do share though if you manage to see some gains.\n\n```python\naugment = Compose([\n            AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.2),\n            TimeStretch(min_rate=0.8, max_rate=1.25, p=0.2),\n            PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n            Shift(p=0.2),\n            Trim(top_db=30.0, p=1.0)\n        ])\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2467510,
      "author_name": "bhavesh1335",
      "author_url": "",
      "post_date": "10/04/2023 16:05:06",
      "content": "<p>Data augmentation can indeed be beneficial for audio data, introducing variations to improve model robustness. However, it's crucial to strike a balance, as excessive augmentation may introduce artifacts. If you've had success or insights, please consider upvoting to share your experiences and foster discussions on effective audio data augmentation strategies. Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2473015,
      "author_name": "leonardoboulitreau",
      "author_url": "",
      "post_date": "10/07/2023 19:31:06",
      "content": "<p>May I ask, did you get any gains with this approach?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2473282,
          "author_name": "msthil",
          "author_url": "",
          "post_date": "10/08/2023 06:34:57",
          "content": "<p>Unfortunately, I haven't had the time to thoroughly try this, however, I did do a quick and dirt training of about 1000 steps using the below augmentation parameters and the loss decreased out from (what I assume was) a local minimum I seemed to have been caught in. That decrease though wasn't too crazy. Hoping to have more time over the next few days to go crazy on a bigger GPU. Please do share though if you manage to see some gains.</p>\n<pre><code>augment = Compose([\n            AddGaussianNoise(min_amplitude=, max_amplitude=, p=),\n            TimeStretch(min_rate=, max_rate=, p=),\n            PitchShift(min_semitones=-, max_semitones=, p=),\n            Shift(p=),\n            Trim(top_db=, p=)\n        ])\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2467447": "Has anyone had any luck training with data augmentation strategies? I know data augmentation is really helpful in CV (unless some new research paper I haven't read has overturned that), but I was surprised (maybe due to my naivety in DL audio) to find out audio had so many augmentation types.\n\nI slapped together a notebook, basically replicating [The Sound of AI's](https://www.youtube.com/playlist?list=PL-wATfeyAMNoR4aqS-Fv0GRmS6bx5RtTW) tutorial. For a sample on LibriSpeech, seems the distortions aren't too bad, but there are enough artifacts for me to question efficacy. CTC loss seems to decrease when using them, but I haven't trained on a mass scale.\n\nMaybe someone who's a native speaker can attest to the quality of the Bengali audio when transformed?\n\n[https://www.kaggle.com/code/msthil/bengali-data-augmentation ](url)",
    "2467510": "Data augmentation can indeed be beneficial for audio data, introducing variations to improve model robustness. However, it's crucial to strike a balance, as excessive augmentation may introduce artifacts. If you've had success or insights, please consider upvoting to share your experiences and foster discussions on effective audio data augmentation strategies. Thank you!",
    "2473015": "May I ask, did you get any gains with this approach?",
    "2473282": "Unfortunately, I haven't had the time to thoroughly try this, however, I did do a quick and dirt training of about 1000 steps using the below augmentation parameters and the loss decreased out from (what I assume was) a local minimum I seemed to have been caught in. That decrease though wasn't too crazy. Hoping to have more time over the next few days to go crazy on a bigger GPU. Please do share though if you manage to see some gains.\n\n```python\naugment = Compose([\n            AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.2),\n            TimeStretch(min_rate=0.8, max_rate=1.25, p=0.2),\n            PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n            Shift(p=0.2),\n            Trim(top_db=30.0, p=1.0)\n        ])\n```"
  },
  "source": "meta"
}