{
  "id": 90464,
  "title": "SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/90464",
  "author_name": "",
  "post_date": "2019-04-24T03:07:57.064082400Z",
  "votes": 34,
  "comment_count": 10,
  "views": 0,
  "content": "<p><a href=\"https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html\">https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html</a></p>\n\n<p>As many people is talking about this paper in twitter posts, I'm just sharing this.\nWhat is proposed in the paper would be simply applicable to this competition.</p>\n\n<p>My notes:\n- <em>'tend to overfit the training data and have a hard time generalizing'</em> -&gt; What we actually do right now.\n- <em>'we take a new approach to augmenting audio data, treating it as a visual problem rather than an audio one'</em> -&gt; Image augmentation would simply work. (and we've already done...)\n- <em>'SpecAugment modifies the spectrogram by warping it in the time direction,...'</em> -&gt; Sounds normal image augmentation in the time direction.\n- <em>'...masking blocks of consecutive frequency channels, and masking blocks of utterances in time.'</em> -&gt; Random erasing (<a href=\"https://arxiv.org/pdf/1708.04896.pdf\">https://arxiv.org/pdf/1708.04896.pdf</a>) would do the similar effect.</p>",
  "messages": [
    {
      "id": "522208",
      "postDate": "04/24/2019 03:07:57",
      "content": "<p><a href=\"https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html\">https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html</a></p>\n\n<p>As many people is talking about this paper in twitter posts, I'm just sharing this.\nWhat is proposed in the paper would be simply applicable to this competition.</p>\n\n<p>My notes:\n- <em>'tend to overfit the training data and have a hard time generalizing'</em> -&gt; What we actually do right now.\n- <em>'we take a new approach to augmenting audio data, treating it as a visual problem rather than an audio one'</em> -&gt; Image augmentation would simply work. (and we've already done...)\n- <em>'SpecAugment modifies the spectrogram by warping it in the time direction,...'</em> -&gt; Sounds normal image augmentation in the time direction.\n- <em>'...masking blocks of consecutive frequency channels, and masking blocks of utterances in time.'</em> -&gt; Random erasing (<a href=\"https://arxiv.org/pdf/1708.04896.pdf\">https://arxiv.org/pdf/1708.04896.pdf</a>) would do the similar effect.</p>",
      "rawMarkdown": "https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html\n\nAs many people is talking about this paper in twitter posts, I'm just sharing this.\nWhat is proposed in the paper would be simply applicable to this competition.\n\nMy notes:\n- _'tend to overfit the training data and have a hard time generalizing'_ -&gt; What we actually do right now.\n- _'we take a new approach to augmenting audio data, treating it as a visual problem rather than an audio one'_ -&gt; Image augmentation would simply work. (and we've already done...)\n- _'SpecAugment modifies the spectrogram by warping it in the time direction,...'_ -&gt; Sounds normal image augmentation in the time direction.\n- _'...masking blocks of consecutive frequency channels, and masking blocks of utterances in time.'_ -&gt; Random erasing (https://arxiv.org/pdf/1708.04896.pdf) would do the similar effect.",
      "votes": null
    },
    {
      "id": "522802",
      "postDate": "04/25/2019 03:49:18",
      "content": "<p>Can you please suggest how can we warp the spectrograms in time direction.? </p>",
      "rawMarkdown": "Can you please suggest how can we warp the spectrograms in time direction.?",
      "votes": null
    },
    {
      "id": "522809",
      "postDate": "04/25/2019 04:19:02",
      "content": "<p>Hey have you read the paper? You will get answer if you do:\n<em>'Time warping is applied via the function <code>sparse_image_warp</code> of tensorflow'</em></p>",
      "rawMarkdown": "Hey have you read the paper? You will get answer if you do:\n_'Time warping is applied via the function `sparse_image_warp` of tensorflow'_",
      "votes": null
    },
    {
      "id": "522852",
      "postDate": "04/25/2019 06:20:52",
      "content": "<p>PyTorch implementation is already released. I'll try it later.\n<a href=\"https://github.com/zcaceres/spec_augment\">https://github.com/zcaceres/spec_augment</a></p>",
      "rawMarkdown": "PyTorch implementation is already released. I'll try it later.\nhttps://github.com/zcaceres/spec_augment",
      "votes": null
    },
    {
      "id": "522927",
      "postDate": "04/25/2019 08:41:16",
      "content": "<p>thanks alot <a href=\"/mhiro2\">@mhiro2</a> </p>",
      "rawMarkdown": "thanks alot @mhiro2",
      "votes": null
    },
    {
      "id": "523011",
      "postDate": "04/25/2019 11:35:24",
      "content": "<p>Thanks! It's quick implementation...</p>",
      "rawMarkdown": "Thanks! It's quick implementation...",
      "votes": null
    },
    {
      "id": "523089",
      "postDate": "04/25/2019 15:00:12",
      "content": "<p>This is how I'm doing it (without warping). Dims are <em>(time, feats)</em>, works on percentages <em>(0,1)</em> that means: for *num_mask* times replace a random number of frames that's between <em>0-percentage</em> of all example frames.</p>\n\n<p>Works with numpy on a single spectrogram</p>\n\n<p>```\ndef spec_augment(spec, num_mask=2, freq_masking_max_percentage=0.3, time_masking_max_percentage=0.3):\n    spec = spec.copy()\n    for i in range(num_mask):\n        freq_percentage = random.uniform(0.0, freq_masking_max_percentage)</p>\n\n<pre><code>    num_freqs_to_mask = int(round(freq_percentage * spec.shape[1]))\n    f0 = np.random.uniform(low=0.0, high=spec.shape[1] - num_freqs_to_mask)\n    f0 = int(f0)\n    spec[:, f0:f0 + num_freqs_to_mask] = 0\n\n    time_percentage = random.uniform(0.0, freq_masking_max_percentage)\n\n    num_frames_to_mask = int(round(time_percentage * spec.shape[0]))\n    t0 = np.random.uniform(low=0.0, high=spec.shape[0] - num_frames_to_mask)\n    t0 = int(t0)\n    spec[t0:t0 + num_frames_to_mask, :] = 0\n\nreturn spec\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "This is how I'm doing it (without warping). Dims are *(time, feats)*, works on percentages *(0,1)* that means: for *num_mask* times replace a random number of frames that's between *0-percentage* of all example frames.\n\nWorks with numpy on a single spectrogram\n\n```\ndef spec_augment(spec, num_mask=2, freq_masking_max_percentage=0.3, time_masking_max_percentage=0.3):\n    spec = spec.copy()\n    for i in range(num_mask):\n        freq_percentage = random.uniform(0.0, freq_masking_max_percentage)\n        \n        num_freqs_to_mask = int(round(freq_percentage * spec.shape[1]))\n        f0 = np.random.uniform(low=0.0, high=spec.shape[1] - num_freqs_to_mask)\n        f0 = int(f0)\n        spec[:, f0:f0 + num_freqs_to_mask] = 0\n\n        time_percentage = random.uniform(0.0, freq_masking_max_percentage)\n        \n        num_frames_to_mask = int(round(time_percentage * spec.shape[0]))\n        t0 = np.random.uniform(low=0.0, high=spec.shape[0] - num_frames_to_mask)\n        t0 = int(t0)\n        spec[t0:t0 + num_frames_to_mask, :] = 0\n    \n    return spec\n    \n```",
      "votes": null
    },
    {
      "id": "523106",
      "postDate": "04/25/2019 15:23:40",
      "content": "<p>Actually, I made kernel with an example:\n<a href=\"https://www.kaggle.com/davids1992/specaugment-quick-implementation/\">https://www.kaggle.com/davids1992/specaugment-quick-implementation/</a></p>",
      "rawMarkdown": "Actually, I made kernel with an example:\nhttps://www.kaggle.com/davids1992/specaugment-quick-implementation/",
      "votes": null
    },
    {
      "id": "523788",
      "postDate": "04/27/2019 02:44:07",
      "content": "<p>I just gave it a shot. It increased my CV around 0.02-0.03, but LB stayed the same.</p>",
      "rawMarkdown": "I just gave it a shot. It increased my CV around 0.02-0.03, but LB stayed the same.",
      "votes": null
    },
    {
      "id": "523851",
      "postDate": "04/27/2019 08:11:21",
      "content": "<p>My suspicion: SpecAugment is robust in speech recognition, where there are 2 networks: encoder, for acoustic signal, and decoder,  that works more on language. Maybe SpecAugment surprisingly has a stronger effect on decoder network because masking some acoustic frames doesn't help nor disturb encoder mechanics, but decoder tends to predict some common language phenomena even if they don't occur in given encoder representation at some moment.</p>",
      "rawMarkdown": "My suspicion: SpecAugment is robust in speech recognition, where there are 2 networks: encoder, for acoustic signal, and decoder,  that works more on language. Maybe SpecAugment surprisingly has a stronger effect on decoder network because masking some acoustic frames doesn't help nor disturb encoder mechanics, but decoder tends to predict some common language phenomena even if they don't occur in given encoder representation at some moment.",
      "votes": null
    },
    {
      "id": "523852",
      "postDate": "04/27/2019 08:11:58",
      "content": "<p>Or: setting percentage to some value results in masking relatively more frames for longer examples than for short examples, and maybe it's too many?</p>",
      "rawMarkdown": "Or: setting percentage to some value results in masking relatively more frames for longer examples than for short examples, and maybe it's too many?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 522802,
      "author_name": "kbhartiya83",
      "author_url": "",
      "post_date": "04/25/2019 03:49:18",
      "content": "<p>Can you please suggest how can we warp the spectrograms in time direction.? </p>",
      "votes": null,
      "replies": [
        {
          "id": 522809,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "04/25/2019 04:19:02",
          "content": "<p>Hey have you read the paper? You will get answer if you do:\n<em>'Time warping is applied via the function <code>sparse_image_warp</code> of tensorflow'</em></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522852,
      "author_name": "mhiro2",
      "author_url": "",
      "post_date": "04/25/2019 06:20:52",
      "content": "<p>PyTorch implementation is already released. I'll try it later.\n<a href=\"https://github.com/zcaceres/spec_augment\">https://github.com/zcaceres/spec_augment</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 522927,
          "author_name": "kbhartiya83",
          "author_url": "",
          "post_date": "04/25/2019 08:41:16",
          "content": "<p>thanks alot <a href=\"/mhiro2\">@mhiro2</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523011,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "04/25/2019 11:35:24",
          "content": "<p>Thanks! It's quick implementation...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523089,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "04/25/2019 15:00:12",
      "content": "<p>This is how I'm doing it (without warping). Dims are <em>(time, feats)</em>, works on percentages <em>(0,1)</em> that means: for *num_mask* times replace a random number of frames that's between <em>0-percentage</em> of all example frames.</p>\n\n<p>Works with numpy on a single spectrogram</p>\n\n<p>```\ndef spec_augment(spec, num_mask=2, freq_masking_max_percentage=0.3, time_masking_max_percentage=0.3):\n    spec = spec.copy()\n    for i in range(num_mask):\n        freq_percentage = random.uniform(0.0, freq_masking_max_percentage)</p>\n\n<pre><code>    num_freqs_to_mask = int(round(freq_percentage * spec.shape[1]))\n    f0 = np.random.uniform(low=0.0, high=spec.shape[1] - num_freqs_to_mask)\n    f0 = int(f0)\n    spec[:, f0:f0 + num_freqs_to_mask] = 0\n\n    time_percentage = random.uniform(0.0, freq_masking_max_percentage)\n\n    num_frames_to_mask = int(round(time_percentage * spec.shape[0]))\n    t0 = np.random.uniform(low=0.0, high=spec.shape[0] - num_frames_to_mask)\n    t0 = int(t0)\n    spec[t0:t0 + num_frames_to_mask, :] = 0\n\nreturn spec\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 523106,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/25/2019 15:23:40",
          "content": "<p>Actually, I made kernel with an example:\n<a href=\"https://www.kaggle.com/davids1992/specaugment-quick-implementation/\">https://www.kaggle.com/davids1992/specaugment-quick-implementation/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523788,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "04/27/2019 02:44:07",
          "content": "<p>I just gave it a shot. It increased my CV around 0.02-0.03, but LB stayed the same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523851,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/27/2019 08:11:21",
          "content": "<p>My suspicion: SpecAugment is robust in speech recognition, where there are 2 networks: encoder, for acoustic signal, and decoder,  that works more on language. Maybe SpecAugment surprisingly has a stronger effect on decoder network because masking some acoustic frames doesn't help nor disturb encoder mechanics, but decoder tends to predict some common language phenomena even if they don't occur in given encoder representation at some moment.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523852,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/27/2019 08:11:58",
          "content": "<p>Or: setting percentage to some value results in masking relatively more frames for longer examples than for short examples, and maybe it's too many?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522208": "https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html\n\nAs many people is talking about this paper in twitter posts, I'm just sharing this.\nWhat is proposed in the paper would be simply applicable to this competition.\n\nMy notes:\n- _'tend to overfit the training data and have a hard time generalizing'_ -&gt; What we actually do right now.\n- _'we take a new approach to augmenting audio data, treating it as a visual problem rather than an audio one'_ -&gt; Image augmentation would simply work. (and we've already done...)\n- _'SpecAugment modifies the spectrogram by warping it in the time direction,...'_ -&gt; Sounds normal image augmentation in the time direction.\n- _'...masking blocks of consecutive frequency channels, and masking blocks of utterances in time.'_ -&gt; Random erasing (https://arxiv.org/pdf/1708.04896.pdf) would do the similar effect.",
    "522802": "Can you please suggest how can we warp the spectrograms in time direction.?",
    "522809": "Hey have you read the paper? You will get answer if you do:\n_'Time warping is applied via the function `sparse_image_warp` of tensorflow'_",
    "522852": "PyTorch implementation is already released. I'll try it later.\nhttps://github.com/zcaceres/spec_augment",
    "522927": "thanks alot @mhiro2",
    "523011": "Thanks! It's quick implementation...",
    "523089": "This is how I'm doing it (without warping). Dims are *(time, feats)*, works on percentages *(0,1)* that means: for *num_mask* times replace a random number of frames that's between *0-percentage* of all example frames.\n\nWorks with numpy on a single spectrogram\n\n```\ndef spec_augment(spec, num_mask=2, freq_masking_max_percentage=0.3, time_masking_max_percentage=0.3):\n    spec = spec.copy()\n    for i in range(num_mask):\n        freq_percentage = random.uniform(0.0, freq_masking_max_percentage)\n        \n        num_freqs_to_mask = int(round(freq_percentage * spec.shape[1]))\n        f0 = np.random.uniform(low=0.0, high=spec.shape[1] - num_freqs_to_mask)\n        f0 = int(f0)\n        spec[:, f0:f0 + num_freqs_to_mask] = 0\n\n        time_percentage = random.uniform(0.0, freq_masking_max_percentage)\n        \n        num_frames_to_mask = int(round(time_percentage * spec.shape[0]))\n        t0 = np.random.uniform(low=0.0, high=spec.shape[0] - num_frames_to_mask)\n        t0 = int(t0)\n        spec[t0:t0 + num_frames_to_mask, :] = 0\n    \n    return spec\n    \n```",
    "523106": "Actually, I made kernel with an example:\nhttps://www.kaggle.com/davids1992/specaugment-quick-implementation/",
    "523788": "I just gave it a shot. It increased my CV around 0.02-0.03, but LB stayed the same.",
    "523851": "My suspicion: SpecAugment is robust in speech recognition, where there are 2 networks: encoder, for acoustic signal, and decoder,  that works more on language. Maybe SpecAugment surprisingly has a stronger effect on decoder network because masking some acoustic frames doesn't help nor disturb encoder mechanics, but decoder tends to predict some common language phenomena even if they don't occur in given encoder representation at some moment.",
    "523852": "Or: setting percentage to some value results in masking relatively more frames for longer examples than for short examples, and maybe it's too many?"
  },
  "source": "meta"
}