{
  "id": 174494,
  "title": "Augmentation Methods",
  "url": "/competitions/birdsong-recognition/discussion/174494",
  "author_name": "",
  "post_date": "2020-08-13T19:06:37.393850700Z",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In additional to topic <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170913\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/170913</a> with great visualizations I want to add some augmentation ideas with corresponding implementations.</p>\n<p><strong>Audiomentations</strong></p>\n<p>Great audio implementations in one of github repositories (link below): you can see about 20 augmentation techniques with short descriptions in front of each class. Some of those:</p>\n<ul>\n<li><em>AddImpulseResponse, AddGaussianNoise, AddBackgroundNoise, AddShortNoises</em>.</li>\n<li><em>TimeMask, FrequencyMask</em>.</li>\n<li><em>Shift, PitchShift, TimeStretch</em>.</li>\n</ul>\n<p>code: <a href=\"https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\" target=\"_blank\">https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py</a></p>\n<p><strong>SpecAugment</strong></p>\n<p>SpecAugment consists of three consecutive augmenatation methods:</p>\n<ul>\n<li><em>Time warp</em>:  exchange some tuple of points in a row audio signal.</li>\n<li><em>Frequency masking</em>: mask some horizontal line inside a mel-spectrogram.</li>\n<li><em>Time masking</em>: mask some vertical line inside a mel-spectrogram.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7a00e5b29b6ecc1c3e56da9b7cbd3012%2FSpecAugmentations.png?generation=1597342806077272&amp;alt=media\" alt=\"\"></p>\n<p>paper: <a href=\"https://arxiv.org/pdf/1904.08779.pdf\" target=\"_blank\">https://arxiv.org/pdf/1904.08779.pdf</a>.<br>\ncode: we've already had an implementation of this method inside one of a public kernels - <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection#Train-SED-model-with-only-weak-supervision\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection#Train-SED-model-with-only-weak-supervision</a> (another thanks for great kernel!).</p>\n<p><strong>MixAugment</strong></p>\n<p>The main contribution of this paper is an idea to use more than one sample inside an augmentation technique (but more than two ones don't increase a target metrics). For two samples <code>(x1, y1)</code> and <code>(x2, y2)</code> we can do it next way:</p>\n<pre><code>x = lambda * x1  + (1 - lambda) * x2\ny = lambda * y1  + (1 - lambda) * y2\n</code></pre>\n<p>It called <em>mixups</em>, but there are lots of other techniques of augmentations including two samples (see the paper for a details):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc1510aa28daabe7651d42eed02a2cfe%2FMixAugment.png?generation=1597344047630223&amp;alt=media\" alt=\"\"></p>\n<p>paper: <a href=\"https://arxiv.org/pdf/1805.11272.pdf\" target=\"_blank\">https://arxiv.org/pdf/1805.11272.pdf</a>. <br>\ncode: <a href=\"https://github.com/ceciliaresearch/MixedExample/blob/master/mixed_example.py\" target=\"_blank\">https://github.com/ceciliaresearch/MixedExample/blob/master/mixed_example.py</a></p>\n<p><strong>How change the final set of augmentations?</strong></p>\n<p>In a last paper autors just applyed each augmentation separetly and saw on results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F57e2abe4418ca8046f0c411fdf00b572%2FAugmentResults.png?generation=1597345065735225&amp;alt=media\" alt=\"\"></p>\n<p>Top of these methods can be chosen as a final set of augmentation methods and can be applied to the train set together.</p>",
  "messages": [
    {
      "id": "969548",
      "postDate": "08/13/2020 19:06:37",
      "content": "<p>In additional to topic <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170913\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/170913</a> with great visualizations I want to add some augmentation ideas with corresponding implementations.</p>\n<p><strong>Audiomentations</strong></p>\n<p>Great audio implementations in one of github repositories (link below): you can see about 20 augmentation techniques with short descriptions in front of each class. Some of those:</p>\n<ul>\n<li><em>AddImpulseResponse, AddGaussianNoise, AddBackgroundNoise, AddShortNoises</em>.</li>\n<li><em>TimeMask, FrequencyMask</em>.</li>\n<li><em>Shift, PitchShift, TimeStretch</em>.</li>\n</ul>\n<p>code: <a href=\"https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\" target=\"_blank\">https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py</a></p>\n<p><strong>SpecAugment</strong></p>\n<p>SpecAugment consists of three consecutive augmenatation methods:</p>\n<ul>\n<li><em>Time warp</em>:  exchange some tuple of points in a row audio signal.</li>\n<li><em>Frequency masking</em>: mask some horizontal line inside a mel-spectrogram.</li>\n<li><em>Time masking</em>: mask some vertical line inside a mel-spectrogram.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7a00e5b29b6ecc1c3e56da9b7cbd3012%2FSpecAugmentations.png?generation=1597342806077272&amp;alt=media\" alt=\"\"></p>\n<p>paper: <a href=\"https://arxiv.org/pdf/1904.08779.pdf\" target=\"_blank\">https://arxiv.org/pdf/1904.08779.pdf</a>.<br>\ncode: we've already had an implementation of this method inside one of a public kernels - <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection#Train-SED-model-with-only-weak-supervision\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection#Train-SED-model-with-only-weak-supervision</a> (another thanks for great kernel!).</p>\n<p><strong>MixAugment</strong></p>\n<p>The main contribution of this paper is an idea to use more than one sample inside an augmentation technique (but more than two ones don't increase a target metrics). For two samples <code>(x1, y1)</code> and <code>(x2, y2)</code> we can do it next way:</p>\n<pre><code>x = lambda * x1  + (1 - lambda) * x2\ny = lambda * y1  + (1 - lambda) * y2\n</code></pre>\n<p>It called <em>mixups</em>, but there are lots of other techniques of augmentations including two samples (see the paper for a details):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc1510aa28daabe7651d42eed02a2cfe%2FMixAugment.png?generation=1597344047630223&amp;alt=media\" alt=\"\"></p>\n<p>paper: <a href=\"https://arxiv.org/pdf/1805.11272.pdf\" target=\"_blank\">https://arxiv.org/pdf/1805.11272.pdf</a>. <br>\ncode: <a href=\"https://github.com/ceciliaresearch/MixedExample/blob/master/mixed_example.py\" target=\"_blank\">https://github.com/ceciliaresearch/MixedExample/blob/master/mixed_example.py</a></p>\n<p><strong>How change the final set of augmentations?</strong></p>\n<p>In a last paper autors just applyed each augmentation separetly and saw on results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F57e2abe4418ca8046f0c411fdf00b572%2FAugmentResults.png?generation=1597345065735225&amp;alt=media\" alt=\"\"></p>\n<p>Top of these methods can be chosen as a final set of augmentation methods and can be applied to the train set together.</p>",
      "rawMarkdown": "In additional to topic https://www.kaggle.com/c/birdsong-recognition/discussion/170913 with great visualizations I want to add some augmentation ideas with corresponding implementations.\n\n**Audiomentations**\n\nGreat audio implementations in one of github repositories (link below): you can see about 20 augmentation techniques with short descriptions in front of each class. Some of those:\n  - *AddImpulseResponse, AddGaussianNoise, AddBackgroundNoise, AddShortNoises*.\n  - *TimeMask, FrequencyMask*.\n  - *Shift, PitchShift, TimeStretch*.\n\ncode: https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\n\n**SpecAugment**\n\n SpecAugment consists of three consecutive augmenatation methods:\n\n  - *Time warp*:  exchange some tuple of points in a row audio signal.\n  - *Frequency masking*: mask some horizontal line inside a mel-spectrogram.\n  - *Time masking*: mask some vertical line inside a mel-spectrogram.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7a00e5b29b6ecc1c3e56da9b7cbd3012%2FSpecAugmentations.png?generation=1597342806077272&alt=media)\n\npaper: https://arxiv.org/pdf/1904.08779.pdf.\ncode: we've already had an implementation of this method inside one of a public kernels - https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection#Train-SED-model-with-only-weak-supervision (another thanks for great kernel!).\n\n**MixAugment**\n\nThe main contribution of this paper is an idea to use more than one sample inside an augmentation technique (but more than two ones don't increase a target metrics). For two samples `(x1, y1)` and `(x2, y2)` we can do it next way:\n\n\n```\nx = lambda * x1  + (1 - lambda) * x2\ny = lambda * y1  + (1 - lambda) * y2\n```\n\nIt called *mixups*, but there are lots of other techniques of augmentations including two samples (see the paper for a details):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc1510aa28daabe7651d42eed02a2cfe%2FMixAugment.png?generation=1597344047630223&alt=media)\n\npaper: https://arxiv.org/pdf/1805.11272.pdf. \ncode: https://github.com/ceciliaresearch/MixedExample/blob/master/mixed_example.py\n\n**How change the final set of augmentations?**\n\nIn a last paper autors just applyed each augmentation separetly and saw on results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F57e2abe4418ca8046f0c411fdf00b572%2FAugmentResults.png?generation=1597345065735225&alt=media)\n\nTop of these methods can be chosen as a final set of augmentation methods and can be applied to the train set together.",
      "votes": null
    },
    {
      "id": "970657",
      "postDate": "08/14/2020 16:37:56",
      "content": "<p>Great resources. Learned some new techniques of augmentations. </p>",
      "rawMarkdown": "Great resources. Learned some new techniques of augmentations.",
      "votes": null
    },
    {
      "id": "970759",
      "postDate": "08/14/2020 18:41:57",
      "content": "<p>You're welcome! </p>",
      "rawMarkdown": "You're welcome!",
      "votes": null
    },
    {
      "id": "971359",
      "postDate": "08/15/2020 12:56:04",
      "content": "<p>Thanks for sharing, great work Sir.</p>",
      "rawMarkdown": "Thanks for sharing, great work Sir.",
      "votes": null
    },
    {
      "id": "971503",
      "postDate": "08/15/2020 15:42:35",
      "content": "<p>Thank you! </p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1001826",
      "postDate": "09/07/2020 15:47:51",
      "content": "<p>thank you for sharing, useful topic 🙏</p>",
      "rawMarkdown": "thank you for sharing, useful topic 🙏",
      "votes": null
    },
    {
      "id": "1001846",
      "postDate": "09/07/2020 16:04:51",
      "content": "<p>You're welcome! </p>",
      "rawMarkdown": "You're welcome!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 971359,
      "author_name": "palaksood97",
      "author_url": "",
      "post_date": "08/15/2020 12:56:04",
      "content": "<p>Thanks for sharing, great work Sir.</p>",
      "votes": null,
      "replies": [
        {
          "id": 971503,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/15/2020 15:42:35",
          "content": "<p>Thank you! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1001826,
      "author_name": "",
      "author_url": "",
      "post_date": "09/07/2020 15:47:51",
      "content": "<p>thank you for sharing, useful topic 🙏</p>",
      "votes": null,
      "replies": [
        {
          "id": 1001846,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/07/2020 16:04:51",
          "content": "<p>You're welcome! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970657,
      "author_name": "azaemon",
      "author_url": "",
      "post_date": "08/14/2020 16:37:56",
      "content": "<p>Great resources. Learned some new techniques of augmentations. </p>",
      "votes": null,
      "replies": [
        {
          "id": 970759,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/14/2020 18:41:57",
          "content": "<p>You're welcome! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "969548": "In additional to topic https://www.kaggle.com/c/birdsong-recognition/discussion/170913 with great visualizations I want to add some augmentation ideas with corresponding implementations.\n\n**Audiomentations**\n\nGreat audio implementations in one of github repositories (link below): you can see about 20 augmentation techniques with short descriptions in front of each class. Some of those:\n  - *AddImpulseResponse, AddGaussianNoise, AddBackgroundNoise, AddShortNoises*.\n  - *TimeMask, FrequencyMask*.\n  - *Shift, PitchShift, TimeStretch*.\n\ncode: https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\n\n**SpecAugment**\n\n SpecAugment consists of three consecutive augmenatation methods:\n\n  - *Time warp*:  exchange some tuple of points in a row audio signal.\n  - *Frequency masking*: mask some horizontal line inside a mel-spectrogram.\n  - *Time masking*: mask some vertical line inside a mel-spectrogram.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7a00e5b29b6ecc1c3e56da9b7cbd3012%2FSpecAugmentations.png?generation=1597342806077272&alt=media)\n\npaper: https://arxiv.org/pdf/1904.08779.pdf.\ncode: we've already had an implementation of this method inside one of a public kernels - https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection#Train-SED-model-with-only-weak-supervision (another thanks for great kernel!).\n\n**MixAugment**\n\nThe main contribution of this paper is an idea to use more than one sample inside an augmentation technique (but more than two ones don't increase a target metrics). For two samples `(x1, y1)` and `(x2, y2)` we can do it next way:\n\n\n```\nx = lambda * x1  + (1 - lambda) * x2\ny = lambda * y1  + (1 - lambda) * y2\n```\n\nIt called *mixups*, but there are lots of other techniques of augmentations including two samples (see the paper for a details):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc1510aa28daabe7651d42eed02a2cfe%2FMixAugment.png?generation=1597344047630223&alt=media)\n\npaper: https://arxiv.org/pdf/1805.11272.pdf. \ncode: https://github.com/ceciliaresearch/MixedExample/blob/master/mixed_example.py\n\n**How change the final set of augmentations?**\n\nIn a last paper autors just applyed each augmentation separetly and saw on results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F57e2abe4418ca8046f0c411fdf00b572%2FAugmentResults.png?generation=1597345065735225&alt=media)\n\nTop of these methods can be chosen as a final set of augmentation methods and can be applied to the train set together.",
    "970657": "Great resources. Learned some new techniques of augmentations.",
    "970759": "You're welcome!",
    "971359": "Thanks for sharing, great work Sir.",
    "971503": "Thank you!",
    "1001826": "thank you for sharing, useful topic 🙏",
    "1001846": "You're welcome!"
  },
  "source": "meta"
}