{
  "id": 311995,
  "title": "Suggested spectrogram augmentation?",
  "url": "/competitions/birdclef-2022/discussion/311995",
  "author_name": "",
  "post_date": "2022-03-09T22:15:29.729475200Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In the beginning, I would state that I am <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce\" target=\"_blank\">converting the audio to spectrogram images</a> and then performing the task of <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-eda-multi-class-flash-effnet\" target=\"_blank\">multi-label image classification</a>. </p>\n<p>From the training and proper separation of the validation/training per audio file, not just image frames (which may be still leaking as a naive split would put two images from the same audio file to train and validation set) seems that the model was soon overfitting the training data so some regularization and augmentation are needed. I have used the only shift in the time domain but seems insufficient, so any other suggestion?</p>",
  "messages": [
    {
      "id": "1717384",
      "postDate": "03/09/2022 22:15:29",
      "content": "<p>In the beginning, I would state that I am <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce\" target=\"_blank\">converting the audio to spectrogram images</a> and then performing the task of <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-eda-multi-class-flash-effnet\" target=\"_blank\">multi-label image classification</a>. </p>\n<p>From the training and proper separation of the validation/training per audio file, not just image frames (which may be still leaking as a naive split would put two images from the same audio file to train and validation set) seems that the model was soon overfitting the training data so some regularization and augmentation are needed. I have used the only shift in the time domain but seems insufficient, so any other suggestion?</p>",
      "rawMarkdown": "In the beginning, I would state that I am [converting the audio to spectrogram images](https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce) and then performing the task of [multi-label image classification](https://www.kaggle.com/jirkaborovec/birdclef-eda-multi-class-flash-effnet). \n\nFrom the training and proper separation of the validation/training per audio file, not just image frames (which may be still leaking as a naive split would put two images from the same audio file to train and validation set) seems that the model was soon overfitting the training data so some regularization and augmentation are needed. I have used the only shift in the time domain but seems insufficient, so any other suggestion?",
      "votes": null
    },
    {
      "id": "1721322",
      "postDate": "03/13/2022 15:23:29",
      "content": "<p>Try adding Gaussian noise</p>",
      "rawMarkdown": "Try adding Gaussian noise",
      "votes": null
    },
    {
      "id": "1724744",
      "postDate": "03/16/2022 13:26:41",
      "content": "<p>Here's my understanding from what I've seen studying the previous competition. I have provided references to a single paper for all of these augmentations, but I have seen these augmentations mentioned by many and referenced in other papers as well.</p>\n<ul>\n<li>Mixing of images <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Random Power <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>White noise <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Pink Noise <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Bandpass noise <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Lower the upper frequencies <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n</ul>\n<p>People appear to use this library if operating in Pytorch</p>\n<ul>\n<li><a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\"><strong>Audiomentations</strong></a></li>\n</ul>\n<p>For Tensorflow…</p>\n<ul>\n<li>I have noticed that <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/audio\" target=\"_blank\"><strong>tfio.audio</strong></a> has masking audio augmentations for mel-spectrograms (vertical <em>aka time masking</em> and horizontal <em>aka frequency masking</em>). </li>\n<li>You can see this tutorial <a href=\"https://www.tensorflow.org/io/tutorials/audio\" target=\"_blank\"><strong>here</strong></a> for more details. <ul>\n<li>The particular section of interest is <a href=\"https://www.tensorflow.org/io/tutorials/audio#specaugment\" target=\"_blank\"><strong>here</strong></a></li></ul></li>\n</ul>\n<p>I'm sure there are more options than this as well. But I hope this gives you a good starting place.</p>\n<p><br></p>\n<hr>\n<p><strong><em>Addendum for the lazy (like me): The full list of augmentations from <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\"><strong>Audiomentations</strong></a> with descriptions:</em></strong></p>\n<hr>\n<p><strong>AddBackgroundNoise</strong></p>\n<ul>\n<li>Add background noise to the input audio.</li>\n</ul>\n<p><strong>AddColoredNoise</strong></p>\n<ul>\n<li>Add colored noise to the input audio.</li>\n</ul>\n<p><strong>ApplyImpulseResponse</strong></p>\n<ul>\n<li>Convolve the given audio with impulse responses.</li>\n</ul>\n<p><strong>BandPassFilter</strong></p>\n<ul>\n<li>Apply band-pass filtering to the input audio.</li>\n</ul>\n<p><strong>BandStopFilter</strong></p>\n<ul>\n<li>Apply band-stop filtering to the input audio. Also known as notch filter.</li>\n</ul>\n<p><strong>Gain</strong></p>\n<ul>\n<li>Multiply the audio by a random amplitude factor to reduce or increase the volume. This technique can help a model become somewhat invariant to the overall gain of the input audio.</li>\n<li>Warning: This transform can return samples outside the [-1, 1] range, which may lead to clipping or wrap distortion, depending on what you do with the audio in a later stage. See also <a href=\"https://en.wikipedia.org/wiki/Clipping_(audio)#Digital_clipping\" target=\"_blank\">https://en.wikipedia.org/wiki/Clipping_(audio)#Digital_clipping</a></li>\n</ul>\n<p><strong>HighPassFilter</strong></p>\n<ul>\n<li>Apply high-pass filtering to the input audio.</li>\n</ul>\n<p><strong>LowPassFilter</strong></p>\n<ul>\n<li>Apply low-pass filtering to the input audio.</li>\n</ul>\n<p><strong>PeakNormalization</strong></p>\n<ul>\n<li>Apply a constant amount of gain, so that highest signal level present in each audio snippet in the batch becomes 0 dBFS, i.e. the loudest level allowed if all samples must be between -1 and 1.</li>\n<li>This transform has an alternative mode (apply_to=\"only_too_loud_sounds\") where it only applies to audio snippets that have extreme values outside the [-1, 1] range. This is useful for avoiding digital clipping in audio that is too loud, while leaving other audio untouched.</li>\n</ul>\n<p><strong>PitchShift</strong></p>\n<ul>\n<li>Pitch-shift sounds up or down without changing the tempo.</li>\n</ul>\n<p><strong>PolarityInversion</strong></p>\n<ul>\n<li>Flip the audio samples upside-down, reversing their polarity. In other words, multiply the waveform by -1, so negative values become positive, and vice versa. The result will sound the same compared to the original when played back in isolation. However, when mixed with other audio sources, the result may be different. This waveform inversion technique is sometimes used for audio cancellation or obtaining the difference between two waveforms. However, in the context of audio data augmentation, this transform can be useful when training phase-aware machine learning models.</li>\n</ul>\n<p><strong>Shift</strong></p>\n<ul>\n<li>Shift the audio forwards or backwards, with or without rollover</li>\n</ul>\n<p><strong>ShuffleChannels</strong></p>\n<ul>\n<li>Given multichannel audio input (e.g. stereo), shuffle the channels, e.g. so left can become right and vice versa. This transform can help combat positional bias in machine learning models that input multichannel waveforms.</li>\n<li>If the input audio is mono, this transform does nothing except emit a warning.</li>\n</ul>\n<p><strong>TimeInversion</strong></p>\n<ul>\n<li>Reverse (invert) the audio along the time axis similar to random flip of an image in the visual domain. This can be relevant in the context of audio classification. It was successfully applied in the paper AudioCLIP: Extending CLIP to Image, Text and Audio</li>\n</ul>",
      "rawMarkdown": "Here's my understanding from what I've seen studying the previous competition. I have provided references to a single paper for all of these augmentations, but I have seen these augmentations mentioned by many and referenced in other papers as well.\n* Mixing of images [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Random Power [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* White noise [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Pink Noise [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Bandpass noise [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Lower the upper frequencies [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n\nPeople appear to use this library if operating in Pytorch\n* [**Audiomentations**](https://github.com/asteroid-team/torch-audiomentations)\n\nFor Tensorflow...\n* I have noticed that [**tfio.audio**](https://www.tensorflow.org/io/api_docs/python/tfio/audio) has masking audio augmentations for mel-spectrograms (vertical *aka time masking* and horizontal *aka frequency masking*). \n* You can see this tutorial [**here**](https://www.tensorflow.org/io/tutorials/audio) for more details. \n  * The particular section of interest is [**here**](https://www.tensorflow.org/io/tutorials/audio#specaugment)\n\nI'm sure there are more options than this as well. But I hope this gives you a good starting place.\n\n<br>\n\n---\n\n***Addendum for the lazy (like me): The full list of augmentations from [**Audiomentations**](https://github.com/asteroid-team/torch-audiomentations) with descriptions:***\n\n---\n\n**AddBackgroundNoise**\n* Add background noise to the input audio.\n\n**AddColoredNoise**\n* Add colored noise to the input audio.\n\n**ApplyImpulseResponse**\n* Convolve the given audio with impulse responses.\n\n**BandPassFilter**\n* Apply band-pass filtering to the input audio.\n\n**BandStopFilter**\n* Apply band-stop filtering to the input audio. Also known as notch filter.\n\n**Gain**\n* Multiply the audio by a random amplitude factor to reduce or increase the volume. This technique can help a model become somewhat invariant to the overall gain of the input audio.\n* Warning: This transform can return samples outside the [-1, 1] range, which may lead to clipping or wrap distortion, depending on what you do with the audio in a later stage. See also https://en.wikipedia.org/wiki/Clipping_(audio)#Digital_clipping\n\n**HighPassFilter**\n* Apply high-pass filtering to the input audio.\n\n**LowPassFilter**\n* Apply low-pass filtering to the input audio.\n\n**PeakNormalization**\n* Apply a constant amount of gain, so that highest signal level present in each audio snippet in the batch becomes 0 dBFS, i.e. the loudest level allowed if all samples must be between -1 and 1.\n* This transform has an alternative mode (apply_to=\"only_too_loud_sounds\") where it only applies to audio snippets that have extreme values outside the [-1, 1] range. This is useful for avoiding digital clipping in audio that is too loud, while leaving other audio untouched.\n\n**PitchShift**\n* Pitch-shift sounds up or down without changing the tempo.\n\n**PolarityInversion**\n* Flip the audio samples upside-down, reversing their polarity. In other words, multiply the waveform by -1, so negative values become positive, and vice versa. The result will sound the same compared to the original when played back in isolation. However, when mixed with other audio sources, the result may be different. This waveform inversion technique is sometimes used for audio cancellation or obtaining the difference between two waveforms. However, in the context of audio data augmentation, this transform can be useful when training phase-aware machine learning models.\n\n**Shift**\n* Shift the audio forwards or backwards, with or without rollover\n\n**ShuffleChannels**\n* Given multichannel audio input (e.g. stereo), shuffle the channels, e.g. so left can become right and vice versa. This transform can help combat positional bias in machine learning models that input multichannel waveforms.\n* If the input audio is mono, this transform does nothing except emit a warning.\n\n**TimeInversion**\n* Reverse (invert) the audio along the time axis similar to random flip of an image in the visual domain. This can be relevant in the context of audio classification. It was successfully applied in the paper AudioCLIP: Extending CLIP to Image, Text and Audio",
      "votes": null
    },
    {
      "id": "1724764",
      "postDate": "03/16/2022 13:41:31",
      "content": "<p>wow, very helpful inside; thank you</p>",
      "rawMarkdown": "wow, very helpful inside; thank you",
      "votes": null
    },
    {
      "id": "1753438",
      "postDate": "04/12/2022 20:27:53",
      "content": "<p>Did you try Mixup augmentation? Mix up is an augmentation in which a new spectrogram and ground truth by mixing two pairs of data. The detail is described in the paper below.</p>\n<p><a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">mixup: Beyond Empirical Risk Minimization</a></p>",
      "rawMarkdown": "Did you try Mixup augmentation? Mix up is an augmentation in which a new spectrogram and ground truth by mixing two pairs of data. The detail is described in the paper below.\n\n[mixup: Beyond Empirical Risk Minimization](https://arxiv.org/abs/1710.09412)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1721322,
      "author_name": "antorino",
      "author_url": "",
      "post_date": "03/13/2022 15:23:29",
      "content": "<p>Try adding Gaussian noise</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1724744,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "03/16/2022 13:26:41",
      "content": "<p>Here's my understanding from what I've seen studying the previous competition. I have provided references to a single paper for all of these augmentations, but I have seen these augmentations mentioned by many and referenced in other papers as well.</p>\n<ul>\n<li>Mixing of images <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Random Power <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>White noise <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Pink Noise <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Bandpass noise <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n<li>Lower the upper frequencies <a href=\"https://arxiv.org/pdf/2107.04878.pdf\" target=\"_blank\"><em>[REF]</em></a></li>\n</ul>\n<p>People appear to use this library if operating in Pytorch</p>\n<ul>\n<li><a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\"><strong>Audiomentations</strong></a></li>\n</ul>\n<p>For Tensorflow…</p>\n<ul>\n<li>I have noticed that <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/audio\" target=\"_blank\"><strong>tfio.audio</strong></a> has masking audio augmentations for mel-spectrograms (vertical <em>aka time masking</em> and horizontal <em>aka frequency masking</em>). </li>\n<li>You can see this tutorial <a href=\"https://www.tensorflow.org/io/tutorials/audio\" target=\"_blank\"><strong>here</strong></a> for more details. <ul>\n<li>The particular section of interest is <a href=\"https://www.tensorflow.org/io/tutorials/audio#specaugment\" target=\"_blank\"><strong>here</strong></a></li></ul></li>\n</ul>\n<p>I'm sure there are more options than this as well. But I hope this gives you a good starting place.</p>\n<p><br></p>\n<hr>\n<p><strong><em>Addendum for the lazy (like me): The full list of augmentations from <a href=\"https://github.com/asteroid-team/torch-audiomentations\" target=\"_blank\"><strong>Audiomentations</strong></a> with descriptions:</em></strong></p>\n<hr>\n<p><strong>AddBackgroundNoise</strong></p>\n<ul>\n<li>Add background noise to the input audio.</li>\n</ul>\n<p><strong>AddColoredNoise</strong></p>\n<ul>\n<li>Add colored noise to the input audio.</li>\n</ul>\n<p><strong>ApplyImpulseResponse</strong></p>\n<ul>\n<li>Convolve the given audio with impulse responses.</li>\n</ul>\n<p><strong>BandPassFilter</strong></p>\n<ul>\n<li>Apply band-pass filtering to the input audio.</li>\n</ul>\n<p><strong>BandStopFilter</strong></p>\n<ul>\n<li>Apply band-stop filtering to the input audio. Also known as notch filter.</li>\n</ul>\n<p><strong>Gain</strong></p>\n<ul>\n<li>Multiply the audio by a random amplitude factor to reduce or increase the volume. This technique can help a model become somewhat invariant to the overall gain of the input audio.</li>\n<li>Warning: This transform can return samples outside the [-1, 1] range, which may lead to clipping or wrap distortion, depending on what you do with the audio in a later stage. See also <a href=\"https://en.wikipedia.org/wiki/Clipping_(audio)#Digital_clipping\" target=\"_blank\">https://en.wikipedia.org/wiki/Clipping_(audio)#Digital_clipping</a></li>\n</ul>\n<p><strong>HighPassFilter</strong></p>\n<ul>\n<li>Apply high-pass filtering to the input audio.</li>\n</ul>\n<p><strong>LowPassFilter</strong></p>\n<ul>\n<li>Apply low-pass filtering to the input audio.</li>\n</ul>\n<p><strong>PeakNormalization</strong></p>\n<ul>\n<li>Apply a constant amount of gain, so that highest signal level present in each audio snippet in the batch becomes 0 dBFS, i.e. the loudest level allowed if all samples must be between -1 and 1.</li>\n<li>This transform has an alternative mode (apply_to=\"only_too_loud_sounds\") where it only applies to audio snippets that have extreme values outside the [-1, 1] range. This is useful for avoiding digital clipping in audio that is too loud, while leaving other audio untouched.</li>\n</ul>\n<p><strong>PitchShift</strong></p>\n<ul>\n<li>Pitch-shift sounds up or down without changing the tempo.</li>\n</ul>\n<p><strong>PolarityInversion</strong></p>\n<ul>\n<li>Flip the audio samples upside-down, reversing their polarity. In other words, multiply the waveform by -1, so negative values become positive, and vice versa. The result will sound the same compared to the original when played back in isolation. However, when mixed with other audio sources, the result may be different. This waveform inversion technique is sometimes used for audio cancellation or obtaining the difference between two waveforms. However, in the context of audio data augmentation, this transform can be useful when training phase-aware machine learning models.</li>\n</ul>\n<p><strong>Shift</strong></p>\n<ul>\n<li>Shift the audio forwards or backwards, with or without rollover</li>\n</ul>\n<p><strong>ShuffleChannels</strong></p>\n<ul>\n<li>Given multichannel audio input (e.g. stereo), shuffle the channels, e.g. so left can become right and vice versa. This transform can help combat positional bias in machine learning models that input multichannel waveforms.</li>\n<li>If the input audio is mono, this transform does nothing except emit a warning.</li>\n</ul>\n<p><strong>TimeInversion</strong></p>\n<ul>\n<li>Reverse (invert) the audio along the time axis similar to random flip of an image in the visual domain. This can be relevant in the context of audio classification. It was successfully applied in the paper AudioCLIP: Extending CLIP to Image, Text and Audio</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1724764,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/16/2022 13:41:31",
          "content": "<p>wow, very helpful inside; thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1753438,
      "author_name": "kotanoda",
      "author_url": "",
      "post_date": "04/12/2022 20:27:53",
      "content": "<p>Did you try Mixup augmentation? Mix up is an augmentation in which a new spectrogram and ground truth by mixing two pairs of data. The detail is described in the paper below.</p>\n<p><a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">mixup: Beyond Empirical Risk Minimization</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1717384": "In the beginning, I would state that I am [converting the audio to spectrogram images](https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce) and then performing the task of [multi-label image classification](https://www.kaggle.com/jirkaborovec/birdclef-eda-multi-class-flash-effnet). \n\nFrom the training and proper separation of the validation/training per audio file, not just image frames (which may be still leaking as a naive split would put two images from the same audio file to train and validation set) seems that the model was soon overfitting the training data so some regularization and augmentation are needed. I have used the only shift in the time domain but seems insufficient, so any other suggestion?",
    "1721322": "Try adding Gaussian noise",
    "1724744": "Here's my understanding from what I've seen studying the previous competition. I have provided references to a single paper for all of these augmentations, but I have seen these augmentations mentioned by many and referenced in other papers as well.\n* Mixing of images [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Random Power [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* White noise [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Pink Noise [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Bandpass noise [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n* Lower the upper frequencies [*[REF]*](https://arxiv.org/pdf/2107.04878.pdf)\n\nPeople appear to use this library if operating in Pytorch\n* [**Audiomentations**](https://github.com/asteroid-team/torch-audiomentations)\n\nFor Tensorflow...\n* I have noticed that [**tfio.audio**](https://www.tensorflow.org/io/api_docs/python/tfio/audio) has masking audio augmentations for mel-spectrograms (vertical *aka time masking* and horizontal *aka frequency masking*). \n* You can see this tutorial [**here**](https://www.tensorflow.org/io/tutorials/audio) for more details. \n  * The particular section of interest is [**here**](https://www.tensorflow.org/io/tutorials/audio#specaugment)\n\nI'm sure there are more options than this as well. But I hope this gives you a good starting place.\n\n<br>\n\n---\n\n***Addendum for the lazy (like me): The full list of augmentations from [**Audiomentations**](https://github.com/asteroid-team/torch-audiomentations) with descriptions:***\n\n---\n\n**AddBackgroundNoise**\n* Add background noise to the input audio.\n\n**AddColoredNoise**\n* Add colored noise to the input audio.\n\n**ApplyImpulseResponse**\n* Convolve the given audio with impulse responses.\n\n**BandPassFilter**\n* Apply band-pass filtering to the input audio.\n\n**BandStopFilter**\n* Apply band-stop filtering to the input audio. Also known as notch filter.\n\n**Gain**\n* Multiply the audio by a random amplitude factor to reduce or increase the volume. This technique can help a model become somewhat invariant to the overall gain of the input audio.\n* Warning: This transform can return samples outside the [-1, 1] range, which may lead to clipping or wrap distortion, depending on what you do with the audio in a later stage. See also https://en.wikipedia.org/wiki/Clipping_(audio)#Digital_clipping\n\n**HighPassFilter**\n* Apply high-pass filtering to the input audio.\n\n**LowPassFilter**\n* Apply low-pass filtering to the input audio.\n\n**PeakNormalization**\n* Apply a constant amount of gain, so that highest signal level present in each audio snippet in the batch becomes 0 dBFS, i.e. the loudest level allowed if all samples must be between -1 and 1.\n* This transform has an alternative mode (apply_to=\"only_too_loud_sounds\") where it only applies to audio snippets that have extreme values outside the [-1, 1] range. This is useful for avoiding digital clipping in audio that is too loud, while leaving other audio untouched.\n\n**PitchShift**\n* Pitch-shift sounds up or down without changing the tempo.\n\n**PolarityInversion**\n* Flip the audio samples upside-down, reversing their polarity. In other words, multiply the waveform by -1, so negative values become positive, and vice versa. The result will sound the same compared to the original when played back in isolation. However, when mixed with other audio sources, the result may be different. This waveform inversion technique is sometimes used for audio cancellation or obtaining the difference between two waveforms. However, in the context of audio data augmentation, this transform can be useful when training phase-aware machine learning models.\n\n**Shift**\n* Shift the audio forwards or backwards, with or without rollover\n\n**ShuffleChannels**\n* Given multichannel audio input (e.g. stereo), shuffle the channels, e.g. so left can become right and vice versa. This transform can help combat positional bias in machine learning models that input multichannel waveforms.\n* If the input audio is mono, this transform does nothing except emit a warning.\n\n**TimeInversion**\n* Reverse (invert) the audio along the time axis similar to random flip of an image in the visual domain. This can be relevant in the context of audio classification. It was successfully applied in the paper AudioCLIP: Extending CLIP to Image, Text and Audio",
    "1724764": "wow, very helpful inside; thank you",
    "1753438": "Did you try Mixup augmentation? Mix up is an augmentation in which a new spectrogram and ground truth by mixing two pairs of data. The detail is described in the paper below.\n\n[mixup: Beyond Empirical Risk Minimization](https://arxiv.org/abs/1710.09412)"
  },
  "source": "meta"
}