{
  "id": 177930,
  "title": "Difference between image and audio augmentations ",
  "url": "/competitions/birdsong-recognition/discussion/177930",
  "author_name": "",
  "post_date": "2020-08-27T22:25:38.973562100Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>This is my first audio competition so one takes new ground every day.<br>\nThere are some techniques to train and classify, e.g. with and without convert to an audio image like spectrogram.<br>\nSomething that I'm curious about, what's really the difference between classic image/vision and spectrogram augmentation like brightness/contrast,cropping, noice, cutout, stretch etc in terms of the final results, seems that one laborate with the same object? Find no field studies or benchmarks between them when I look around.<br>\nAdding noice, drop block, stretch etc are features in vision/image augmentation as well.<br>\nShould one see it as a complement or should one only use audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.<br>\nClassic image/vision augmentation works quite well in this problem, so just a novice wonder what the differences are to the results, I do understand its diff. to the object.</p>\n<p>Update:<br>\nHere is a good view over the different augmentation techniques<br>\n<a href=\"https://github.com/AgaMiko/data-augmentation-review\" target=\"_blank\">https://github.com/AgaMiko/data-augmentation-review</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9cbbcab816fb87b86603f546144906c7%2Finbox_3924325_58cc4ba00cd89c0d77dc6f5c4c62a744_da_diagram_v2.png?generation=1605863408098189&amp;alt=media\" alt=\"\"></p>\n<p>Evidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.</p>",
  "messages": [
    {
      "id": "988225",
      "postDate": "08/27/2020 22:25:38",
      "content": "<p>This is my first audio competition so one takes new ground every day.<br>\nThere are some techniques to train and classify, e.g. with and without convert to an audio image like spectrogram.<br>\nSomething that I'm curious about, what's really the difference between classic image/vision and spectrogram augmentation like brightness/contrast,cropping, noice, cutout, stretch etc in terms of the final results, seems that one laborate with the same object? Find no field studies or benchmarks between them when I look around.<br>\nAdding noice, drop block, stretch etc are features in vision/image augmentation as well.<br>\nShould one see it as a complement or should one only use audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.<br>\nClassic image/vision augmentation works quite well in this problem, so just a novice wonder what the differences are to the results, I do understand its diff. to the object.</p>\n<p>Update:<br>\nHere is a good view over the different augmentation techniques<br>\n<a href=\"https://github.com/AgaMiko/data-augmentation-review\" target=\"_blank\">https://github.com/AgaMiko/data-augmentation-review</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9cbbcab816fb87b86603f546144906c7%2Finbox_3924325_58cc4ba00cd89c0d77dc6f5c4c62a744_da_diagram_v2.png?generation=1605863408098189&amp;alt=media\" alt=\"\"></p>\n<p>Evidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.</p>",
      "rawMarkdown": "This is my first audio competition so one takes new ground every day.\nThere are some techniques to train and classify, e.g. with and without convert to an audio image like spectrogram.\nSomething that I'm curious about, what's really the difference between classic image/vision and spectrogram augmentation like brightness/contrast,cropping, noice, cutout, stretch etc in terms of the final results, seems that one laborate with the same object? Find no field studies or benchmarks between them when I look around.\nAdding noice, drop block, stretch etc are features in vision/image augmentation as well.\nShould one see it as a complement or should one only use audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.\nClassic image/vision augmentation works quite well in this problem, so just a novice wonder what the differences are to the results, I do understand its diff. to the object.\n\nUpdate:\nHere is a good view over the different augmentation techniques\nhttps://github.com/AgaMiko/data-augmentation-review\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9cbbcab816fb87b86603f546144906c7%2Finbox_3924325_58cc4ba00cd89c0d77dc6f5c4c62a744_da_diagram_v2.png?generation=1605863408098189&alt=media)\n\nEvidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.",
      "votes": null
    },
    {
      "id": "988783",
      "postDate": "08/28/2020 09:18:39",
      "content": "<p>I'm like you, first audio, and discovering stuff as I go.  All I can say for now is that melspectrogram is not  a linear transformation, therefore augmentations on the waveform are not necessarily equivalent to similar image augmentations on the spectrogram.  </p>",
      "rawMarkdown": "I'm like you, first audio, and discovering stuff as I go.  All I can say for now is that melspectrogram is not  a linear transformation, therefore augmentations on the waveform are not necessarily equivalent to similar image augmentations on the spectrogram.",
      "votes": null
    },
    {
      "id": "988785",
      "postDate": "08/28/2020 09:21:40",
      "content": "<p>Not sure if everyone has already read up on this kind of stuff but I thought this was an interesting point from a paper abstract.</p>\n<p><strong>However, spectrogram properties are very different to those of natural images. Instead of an object occupying a contiguous region in a natural image, frequencies of a sound are scattered about the frequency axis of a spectrogram in a pattern unique to that particular sound. Applying conventional convolution neural networks has therefore required extensive hand-tuning, and presented the need to find an architecture better suited to the time–frequency properties of audio.</strong></p>",
      "rawMarkdown": "Not sure if everyone has already read up on this kind of stuff but I thought this was an interesting point from a paper abstract.\n\n**However, spectrogram properties are very different to those of natural images. Instead of an object occupying a contiguous region in a natural image, frequencies of a sound are scattered about the frequency axis of a spectrogram in a pattern unique to that particular sound. Applying conventional convolution neural networks has therefore required extensive hand-tuning, and presented the need to find an architecture better suited to the time–frequency properties of audio.**",
      "votes": null
    },
    {
      "id": "988829",
      "postDate": "08/28/2020 10:10:06",
      "content": "<p>You can use augmentation on both waveforms and spectrograms, but keep in mind that spectrograms have a physical meaning. They are <code>time x frequency</code> representations of your audio.</p>\n<p>Therefore it sometimes does not really make sense to apply some transformations that you would use in computer vision. Here are some of my thought, feel free to correct me.</p>\n<h5>Random Flipping</h5>\n<ul>\n<li>Horizontal flipping plays the audio backward, which is okay but not necessarly useful</li>\n<li>Vertical flipping \"inverts\" the frequencies which does not make sense</li>\n</ul>\n<h5>Random Shifting</h5>\n<ul>\n<li>Shifting vertically shifts the frequencies, which is not something I would advise here since I kinda feel like birds sing at very specific frequencies </li>\n<li>Shifting horizontally is completely fine</li>\n</ul>\n<h5>Random Resizing</h5>\n<ul>\n<li>Resizing on the time axis strecthes the audio which can be okay</li>\n<li>Resizing on the frequency axis is quite a weird thing to do in my opinion</li>\n</ul>\n<h5>Random Rotating</h5>\n<ul>\n<li>It not really have physical meaning but I'm pretty sure it messes up everything</li>\n</ul>\n<h5>Random cropping</h5>\n<ul>\n<li>Cropping on frequency axis : you may lose useful information</li>\n<li>Cropping on the time axis : completely okay</li>\n</ul>",
      "rawMarkdown": "You can use augmentation on both waveforms and spectrograms, but keep in mind that spectrograms have a physical meaning. They are `time x frequency` representations of your audio.\n\nTherefore it sometimes does not really make sense to apply some transformations that you would use in computer vision. Here are some of my thought, feel free to correct me.\n\n##### Random Flipping\n- Horizontal flipping plays the audio backward, which is okay but not necessarly useful\n- Vertical flipping \"inverts\" the frequencies which does not make sense\n\n##### Random Shifting\n- Shifting vertically shifts the frequencies, which is not something I would advise here since I kinda feel like birds sing at very specific frequencies \n- Shifting horizontally is completely fine\n\n##### Random Resizing\n- Resizing on the time axis strecthes the audio which can be okay\n- Resizing on the frequency axis is quite a weird thing to do in my opinion\n\n##### Random Rotating \n- It not really have physical meaning but I'm pretty sure it messes up everything\n\n##### Random cropping\n- Cropping on frequency axis : you may lose useful information\n- Cropping on the time axis : completely okay",
      "votes": null
    },
    {
      "id": "988843",
      "postDate": "08/28/2020 10:22:26",
      "content": "<p>That's exactly one of the questions,<br>\n<code>or should one only use Audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.</code></p>\n<p>Definition of a spectrogram - \"displays signal strength over time at the various frequencies present in a waveform. Spectrograms can be two-dimensional graphs with a third variable represented by color, or three-dimensional graphs with a fourth color variable.\"</p>\n<p>As I see it, there are three augmentation methods to use,<br>\n-- Aug the wave before transform to spectrogram, deforming the audio waveform. Can imagine the computational cost, after just experienced loading and pre-process the wave.<br>\n-- Aug the spectrogram, as described by google research in thier report, SpecAugment.<br>\n-- Aug the image with classic image/vision aug setup.</p>\n<p><strong>What I can't find is a benchmark with one augmentation vs another in audio classification problem, or in combinations together.</strong></p>\n<p>For handle overfittning, besides augmentation, one can also use all kind of techniques like dropout, dropblock, batchcutout, mixup etc All seems to work in this problem.</p>\n<p>What I found so far without using the traditional audio augmentation for wave and spectrogram is that fixed image augmentation is working well, just adding random mixup, batchcutout to the input and random line and erase to the image handles the overfittning significant.</p>\n<p>Then we also have TTA like flip,rotate etc, seem to work in some forms in training but not in inference.</p>\n<p>Maybe there are a perfect mix for training, or vice versa one of them or in combination is reducing the final result. The reason for the question/wonder.</p>\n<p>A important factor is time, not only in the spectrogram dimension, but also in training and inference.<br>\nWhat I shall try besides above is to try speeding up things, use for example torchaudio and run the augmentation within nn.Sequential, see if it helps.</p>",
      "rawMarkdown": "That's exactly one of the questions,\n`or should one only use Audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.`\n\nDefinition of a spectrogram - \"displays signal strength over time at the various frequencies present in a waveform. Spectrograms can be two-dimensional graphs with a third variable represented by color, or three-dimensional graphs with a fourth color variable.\"\n\nAs I see it, there are three augmentation methods to use,\n-- Aug the wave before transform to spectrogram, deforming the audio waveform. Can imagine the computational cost, after just experienced loading and pre-process the wave.\n-- Aug the spectrogram, as described by google research in thier report, SpecAugment.\n-- Aug the image with classic image/vision aug setup.\n\n**What I can't find is a benchmark with one augmentation vs another in audio classification problem, or in combinations together.**\n\nFor handle overfittning, besides augmentation, one can also use all kind of techniques like dropout, dropblock, batchcutout, mixup etc All seems to work in this problem.\n\nWhat I found so far without using the traditional audio augmentation for wave and spectrogram is that fixed image augmentation is working well, just adding random mixup, batchcutout to the input and random line and erase to the image handles the overfittning significant.\n\nThen we also have TTA like flip,rotate etc, seem to work in some forms in training but not in inference.\n\nMaybe there are a perfect mix for training, or vice versa one of them or in combination is reducing the final result. The reason for the question/wonder.\n\nA important factor is time, not only in the spectrogram dimension, but also in training and inference.\nWhat I shall try besides above is to try speeding up things, use for example torchaudio and run the augmentation within nn.Sequential, see if it helps.",
      "votes": null
    },
    {
      "id": "988876",
      "postDate": "08/28/2020 11:10:33",
      "content": "<p>Thanks, yes some of them works in training or not does it worse but none works in inference so I'm more skeptical to non-fixed-image augmentation to the training.</p>",
      "rawMarkdown": "Thanks, yes some of them works in training or not does it worse but none works in inference so I'm more skeptical to non-fixed-image augmentation to the training.",
      "votes": null
    },
    {
      "id": "988880",
      "postDate": "08/28/2020 11:14:43",
      "content": "<p>My biggest task right now is to find an effective way to improve loading in training and inference, best way in training is loading the image directly but makes it more complicated in inference, important with the exact format, all well, fun with new areas! :)</p>",
      "rawMarkdown": "My biggest task right now is to find an effective way to improve loading in training and inference, best way in training is loading the image directly but makes it more complicated in inference, important with the exact format, all well, fun with new areas! :)",
      "votes": null
    },
    {
      "id": "990225",
      "postDate": "08/29/2020 12:55:04",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I'm not sure if we are using the same melspectrogram function, but using the torchaudio implementation (and in accord with mathematical intuition), melspectrogram seems to be a linear transformation? that is, <code>f(a+b) = f(a) + f(b)</code>, where a and b are both signals. Of course if s is a scalar, &gt; &gt; f(s*a) might be equal to <code>s * f(a)</code> or &gt; sqrt(s) * f(a) based on the power you use, but overall the transformation seems to be linear? Happy to be proven wrong though, it's also my first audio comp :)</p>",
      "rawMarkdown": "cpmpml I'm not sure if we are using the same melspectrogram function, but using the torchaudio implementation (and in accord with mathematical intuition), melspectrogram seems to be a linear transformation? that is, `f(a+b) = f(a) + f(b)`, where a and b are both signals. Of course if s is a scalar, > > f(s*a) might be equal to `s * f(a)` or > sqrt(s) * f(a) based on the power you use, but overall the transformation seems to be linear? Happy to be proven wrong though, it's also my first audio comp :)",
      "votes": null
    },
    {
      "id": "990707",
      "postDate": "08/29/2020 19:38:05",
      "content": "<p>librosa implementation is not linear.  It applies a mel scaling to the square of stft.  stft is linear, mel scaling is linear, but the middle squaring is not.</p>\n<p>Edit: torchaudio also uses power=2 as default, it is not linear.</p>",
      "rawMarkdown": "librosa implementation is not linear.  It applies a mel scaling to the square of stft.  stft is linear, mel scaling is linear, but the middle squaring is not.\n\nEdit: torchaudio also uses power=2 as default, it is not linear.",
      "votes": null
    },
    {
      "id": "991083",
      "postDate": "08/30/2020 06:12:12",
      "content": "<p>Please share your findings, I would love to learn more on audios.</p>",
      "rawMarkdown": "Please share your findings, I would love to learn more on audios.",
      "votes": null
    },
    {
      "id": "991805",
      "postDate": "08/30/2020 17:13:11",
      "content": "<p>I also forgot that melspectrogram computation ends with a log transform in librosa.  I guess torchaudio is the same.  </p>\n<p>Definitely not linear.</p>",
      "rawMarkdown": "I also forgot that melspectrogram computation ends with a log transform in librosa.  I guess torchaudio is the same.  \n\nDefinitely not linear.",
      "votes": null
    },
    {
      "id": "992472",
      "postDate": "08/31/2020 07:18:08",
      "content": "<p>Here is a good view over the different augmentation techniques<br>\n<a href=\"https://github.com/AgaMiko/data-augmentation-review\" target=\"_blank\">https://github.com/AgaMiko/data-augmentation-review</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F58cc4ba00cd89c0d77dc6f5c4c62a744%2Fda_diagram_v2.png?generation=1598857683558691&amp;alt=media\" alt=\"\"></p>\n<p>Evidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.</p>",
      "rawMarkdown": "Here is a good view over the different augmentation techniques\nhttps://github.com/AgaMiko/data-augmentation-review\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F58cc4ba00cd89c0d77dc6f5c4c62a744%2Fda_diagram_v2.png?generation=1598857683558691&alt=media)\n\nEvidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.",
      "votes": null
    },
    {
      "id": "992632",
      "postDate": "08/31/2020 09:41:23",
      "content": "<p>Additionally, FFT imposes a trade-off between time and frequency dimensions. One of augmentations can be to shift this trade-off between the two.</p>",
      "rawMarkdown": "Additionally, FFT imposes a trade-off between time and frequency dimensions. One of augmentations can be to shift this trade-off between the two.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 988783,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/28/2020 09:18:39",
      "content": "<p>I'm like you, first audio, and discovering stuff as I go.  All I can say for now is that melspectrogram is not  a linear transformation, therefore augmentations on the waveform are not necessarily equivalent to similar image augmentations on the spectrogram.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 988880,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/28/2020 11:14:43",
          "content": "<p>My biggest task right now is to find an effective way to improve loading in training and inference, best way in training is loading the image directly but makes it more complicated in inference, important with the exact format, all well, fun with new areas! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990225,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "08/29/2020 12:55:04",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I'm not sure if we are using the same melspectrogram function, but using the torchaudio implementation (and in accord with mathematical intuition), melspectrogram seems to be a linear transformation? that is, <code>f(a+b) = f(a) + f(b)</code>, where a and b are both signals. Of course if s is a scalar, &gt; &gt; f(s*a) might be equal to <code>s * f(a)</code> or &gt; sqrt(s) * f(a) based on the power you use, but overall the transformation seems to be linear? Happy to be proven wrong though, it's also my first audio comp :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990707,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/29/2020 19:38:05",
          "content": "<p>librosa implementation is not linear.  It applies a mel scaling to the square of stft.  stft is linear, mel scaling is linear, but the middle squaring is not.</p>\n<p>Edit: torchaudio also uses power=2 as default, it is not linear.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 991083,
          "author_name": "alincijov",
          "author_url": "",
          "post_date": "08/30/2020 06:12:12",
          "content": "<p>Please share your findings, I would love to learn more on audios.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 991805,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/30/2020 17:13:11",
          "content": "<p>I also forgot that melspectrogram computation ends with a log transform in librosa.  I guess torchaudio is the same.  </p>\n<p>Definitely not linear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 988785,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "08/28/2020 09:21:40",
      "content": "<p>Not sure if everyone has already read up on this kind of stuff but I thought this was an interesting point from a paper abstract.</p>\n<p><strong>However, spectrogram properties are very different to those of natural images. Instead of an object occupying a contiguous region in a natural image, frequencies of a sound are scattered about the frequency axis of a spectrogram in a pattern unique to that particular sound. Applying conventional convolution neural networks has therefore required extensive hand-tuning, and presented the need to find an architecture better suited to the time–frequency properties of audio.</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 988829,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "08/28/2020 10:10:06",
      "content": "<p>You can use augmentation on both waveforms and spectrograms, but keep in mind that spectrograms have a physical meaning. They are <code>time x frequency</code> representations of your audio.</p>\n<p>Therefore it sometimes does not really make sense to apply some transformations that you would use in computer vision. Here are some of my thought, feel free to correct me.</p>\n<h5>Random Flipping</h5>\n<ul>\n<li>Horizontal flipping plays the audio backward, which is okay but not necessarly useful</li>\n<li>Vertical flipping \"inverts\" the frequencies which does not make sense</li>\n</ul>\n<h5>Random Shifting</h5>\n<ul>\n<li>Shifting vertically shifts the frequencies, which is not something I would advise here since I kinda feel like birds sing at very specific frequencies </li>\n<li>Shifting horizontally is completely fine</li>\n</ul>\n<h5>Random Resizing</h5>\n<ul>\n<li>Resizing on the time axis strecthes the audio which can be okay</li>\n<li>Resizing on the frequency axis is quite a weird thing to do in my opinion</li>\n</ul>\n<h5>Random Rotating</h5>\n<ul>\n<li>It not really have physical meaning but I'm pretty sure it messes up everything</li>\n</ul>\n<h5>Random cropping</h5>\n<ul>\n<li>Cropping on frequency axis : you may lose useful information</li>\n<li>Cropping on the time axis : completely okay</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 988876,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/28/2020 11:10:33",
          "content": "<p>Thanks, yes some of them works in training or not does it worse but none works in inference so I'm more skeptical to non-fixed-image augmentation to the training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 988843,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "08/28/2020 10:22:26",
      "content": "<p>That's exactly one of the questions,<br>\n<code>or should one only use Audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.</code></p>\n<p>Definition of a spectrogram - \"displays signal strength over time at the various frequencies present in a waveform. Spectrograms can be two-dimensional graphs with a third variable represented by color, or three-dimensional graphs with a fourth color variable.\"</p>\n<p>As I see it, there are three augmentation methods to use,<br>\n-- Aug the wave before transform to spectrogram, deforming the audio waveform. Can imagine the computational cost, after just experienced loading and pre-process the wave.<br>\n-- Aug the spectrogram, as described by google research in thier report, SpecAugment.<br>\n-- Aug the image with classic image/vision aug setup.</p>\n<p><strong>What I can't find is a benchmark with one augmentation vs another in audio classification problem, or in combinations together.</strong></p>\n<p>For handle overfittning, besides augmentation, one can also use all kind of techniques like dropout, dropblock, batchcutout, mixup etc All seems to work in this problem.</p>\n<p>What I found so far without using the traditional audio augmentation for wave and spectrogram is that fixed image augmentation is working well, just adding random mixup, batchcutout to the input and random line and erase to the image handles the overfittning significant.</p>\n<p>Then we also have TTA like flip,rotate etc, seem to work in some forms in training but not in inference.</p>\n<p>Maybe there are a perfect mix for training, or vice versa one of them or in combination is reducing the final result. The reason for the question/wonder.</p>\n<p>A important factor is time, not only in the spectrogram dimension, but also in training and inference.<br>\nWhat I shall try besides above is to try speeding up things, use for example torchaudio and run the augmentation within nn.Sequential, see if it helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992472,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "08/31/2020 07:18:08",
      "content": "<p>Here is a good view over the different augmentation techniques<br>\n<a href=\"https://github.com/AgaMiko/data-augmentation-review\" target=\"_blank\">https://github.com/AgaMiko/data-augmentation-review</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F58cc4ba00cd89c0d77dc6f5c4c62a744%2Fda_diagram_v2.png?generation=1598857683558691&amp;alt=media\" alt=\"\"></p>\n<p>Evidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992632,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "08/31/2020 09:41:23",
      "content": "<p>Additionally, FFT imposes a trade-off between time and frequency dimensions. One of augmentations can be to shift this trade-off between the two.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "988225": "This is my first audio competition so one takes new ground every day.\nThere are some techniques to train and classify, e.g. with and without convert to an audio image like spectrogram.\nSomething that I'm curious about, what's really the difference between classic image/vision and spectrogram augmentation like brightness/contrast,cropping, noice, cutout, stretch etc in terms of the final results, seems that one laborate with the same object? Find no field studies or benchmarks between them when I look around.\nAdding noice, drop block, stretch etc are features in vision/image augmentation as well.\nShould one see it as a complement or should one only use audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.\nClassic image/vision augmentation works quite well in this problem, so just a novice wonder what the differences are to the results, I do understand its diff. to the object.\n\nUpdate:\nHere is a good view over the different augmentation techniques\nhttps://github.com/AgaMiko/data-augmentation-review\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9cbbcab816fb87b86603f546144906c7%2Finbox_3924325_58cc4ba00cd89c0d77dc6f5c4c62a744_da_diagram_v2.png?generation=1605863408098189&alt=media)\n\nEvidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.",
    "988783": "I'm like you, first audio, and discovering stuff as I go.  All I can say for now is that melspectrogram is not  a linear transformation, therefore augmentations on the waveform are not necessarily equivalent to similar image augmentations on the spectrogram.",
    "988785": "Not sure if everyone has already read up on this kind of stuff but I thought this was an interesting point from a paper abstract.\n\n**However, spectrogram properties are very different to those of natural images. Instead of an object occupying a contiguous region in a natural image, frequencies of a sound are scattered about the frequency axis of a spectrogram in a pattern unique to that particular sound. Applying conventional convolution neural networks has therefore required extensive hand-tuning, and presented the need to find an architecture better suited to the time–frequency properties of audio.**",
    "988829": "You can use augmentation on both waveforms and spectrograms, but keep in mind that spectrograms have a physical meaning. They are `time x frequency` representations of your audio.\n\nTherefore it sometimes does not really make sense to apply some transformations that you would use in computer vision. Here are some of my thought, feel free to correct me.\n\n##### Random Flipping\n- Horizontal flipping plays the audio backward, which is okay but not necessarly useful\n- Vertical flipping \"inverts\" the frequencies which does not make sense\n\n##### Random Shifting\n- Shifting vertically shifts the frequencies, which is not something I would advise here since I kinda feel like birds sing at very specific frequencies \n- Shifting horizontally is completely fine\n\n##### Random Resizing\n- Resizing on the time axis strecthes the audio which can be okay\n- Resizing on the frequency axis is quite a weird thing to do in my opinion\n\n##### Random Rotating \n- It not really have physical meaning but I'm pretty sure it messes up everything\n\n##### Random cropping\n- Cropping on frequency axis : you may lose useful information\n- Cropping on the time axis : completely okay",
    "988843": "That's exactly one of the questions,\n`or should one only use Audio aug in this problem/task to not sabotage the spectrogram's unique dimensions.`\n\nDefinition of a spectrogram - \"displays signal strength over time at the various frequencies present in a waveform. Spectrograms can be two-dimensional graphs with a third variable represented by color, or three-dimensional graphs with a fourth color variable.\"\n\nAs I see it, there are three augmentation methods to use,\n-- Aug the wave before transform to spectrogram, deforming the audio waveform. Can imagine the computational cost, after just experienced loading and pre-process the wave.\n-- Aug the spectrogram, as described by google research in thier report, SpecAugment.\n-- Aug the image with classic image/vision aug setup.\n\n**What I can't find is a benchmark with one augmentation vs another in audio classification problem, or in combinations together.**\n\nFor handle overfittning, besides augmentation, one can also use all kind of techniques like dropout, dropblock, batchcutout, mixup etc All seems to work in this problem.\n\nWhat I found so far without using the traditional audio augmentation for wave and spectrogram is that fixed image augmentation is working well, just adding random mixup, batchcutout to the input and random line and erase to the image handles the overfittning significant.\n\nThen we also have TTA like flip,rotate etc, seem to work in some forms in training but not in inference.\n\nMaybe there are a perfect mix for training, or vice versa one of them or in combination is reducing the final result. The reason for the question/wonder.\n\nA important factor is time, not only in the spectrogram dimension, but also in training and inference.\nWhat I shall try besides above is to try speeding up things, use for example torchaudio and run the augmentation within nn.Sequential, see if it helps.",
    "988876": "Thanks, yes some of them works in training or not does it worse but none works in inference so I'm more skeptical to non-fixed-image augmentation to the training.",
    "988880": "My biggest task right now is to find an effective way to improve loading in training and inference, best way in training is loading the image directly but makes it more complicated in inference, important with the exact format, all well, fun with new areas! :)",
    "990225": "cpmpml I'm not sure if we are using the same melspectrogram function, but using the torchaudio implementation (and in accord with mathematical intuition), melspectrogram seems to be a linear transformation? that is, `f(a+b) = f(a) + f(b)`, where a and b are both signals. Of course if s is a scalar, > > f(s*a) might be equal to `s * f(a)` or > sqrt(s) * f(a) based on the power you use, but overall the transformation seems to be linear? Happy to be proven wrong though, it's also my first audio comp :)",
    "990707": "librosa implementation is not linear.  It applies a mel scaling to the square of stft.  stft is linear, mel scaling is linear, but the middle squaring is not.\n\nEdit: torchaudio also uses power=2 as default, it is not linear.",
    "991083": "Please share your findings, I would love to learn more on audios.",
    "991805": "I also forgot that melspectrogram computation ends with a log transform in librosa.  I guess torchaudio is the same.  \n\nDefinitely not linear.",
    "992472": "Here is a good view over the different augmentation techniques\nhttps://github.com/AgaMiko/data-augmentation-review\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F58cc4ba00cd89c0d77dc6f5c4c62a744%2Fda_diagram_v2.png?generation=1598857683558691&alt=media)\n\nEvidently from own testing, vision tta also works well against overfittning in this problem, may be interpreted in many ways.",
    "992632": "Additionally, FFT imposes a trade-off between time and frequency dimensions. One of augmentations can be to shift this trade-off between the two."
  },
  "source": "meta"
}