{
  "id": 47388,
  "title": "pitch_shift data augmentation",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47388",
  "author_name": "",
  "post_date": "2018-01-13T06:26:32.771886400Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Does any one got success with pitch shift data augmentation? I am using librosa to add  pitch shift in my samples but it does not seem to help much I have tried with other data augmentation like mixing noise and time shift that helped me gain 2% LB. In particularly I am using the following code</p>\n\n<pre><code>def pitch_shift(samples, sr=sr):\n    return librosa.effects.pitch_shift(samples, sr=sr, n_steps=random.randint(1, 5))\n</code></pre>\n\n<p>Do I have to change some parameters? \nPercentage of data in training with pitch shift?</p>\n\n<p>edit 1:\nGiven the limited time to do detailed experiments. I have not tried time stretching. Does it help? </p>\n\n<pre><code>librosa.effects.time_stretch(y, rate)\n</code></pre>",
  "messages": [
    {
      "id": "268057",
      "postDate": "01/13/2018 06:26:32",
      "content": "<p>Does any one got success with pitch shift data augmentation? I am using librosa to add  pitch shift in my samples but it does not seem to help much I have tried with other data augmentation like mixing noise and time shift that helped me gain 2% LB. In particularly I am using the following code</p>\n\n<pre><code>def pitch_shift(samples, sr=sr):\n    return librosa.effects.pitch_shift(samples, sr=sr, n_steps=random.randint(1, 5))\n</code></pre>\n\n<p>Do I have to change some parameters? \nPercentage of data in training with pitch shift?</p>\n\n<p>edit 1:\nGiven the limited time to do detailed experiments. I have not tried time stretching. Does it help? </p>\n\n<pre><code>librosa.effects.time_stretch(y, rate)\n</code></pre>",
      "rawMarkdown": "Does any one got success with pitch shift data augmentation? I am using librosa to add  pitch shift in my samples but it does not seem to help much I have tried with other data augmentation like mixing noise and time shift that helped me gain 2% LB. In particularly I am using the following code\n\n    def pitch_shift(samples, sr=sr):\n        return librosa.effects.pitch_shift(samples, sr=sr, n_steps=random.randint(1, 5))\n\nDo I have to change some parameters? \nPercentage of data in training with pitch shift?\n\nedit 1:\nGiven the limited time to do detailed experiments. I have not tried time stretching. Does it help? \n\n    librosa.effects.time_stretch(y, rate)",
      "votes": null
    },
    {
      "id": "268113",
      "postDate": "01/13/2018 11:41:00",
      "content": "<p>For us, the pitch shift hurts LB and is also too slow, so we removed it from our data augmentation pipeline.</p>",
      "rawMarkdown": "For us, the pitch shift hurts LB and is also too slow, so we removed it from our data augmentation pipeline.",
      "votes": null
    },
    {
      "id": "268142",
      "postDate": "01/13/2018 14:17:04",
      "content": "<p>I tried it and also didn't see improvements. When I thought about it a bit more, what would actually be happening is that you just shift the distribution over the frequency bins (assuming you use some kind of spectrogram feature). </p>\n\n<p>So if you had for example 5 frequency bins with the following (imaginary) values [0, 3, 5, 4, 0] then after lowering the pitch you would have [3, 5, 4, 0, 0]. So they all shifted a bin lower.  So from a convolutional perspective this doesn't really change the \"picture\" and as a result doesn't add much variety to the training set. But perhaps I overlooked something.</p>",
      "rawMarkdown": "I tried it and also didn't see improvements. When I thought about it a bit more, what would actually be happening is that you just shift the distribution over the frequency bins (assuming you use some kind of spectrogram feature). \n\nSo if you had for example 5 frequency bins with the following (imaginary) values [0, 3, 5, 4, 0] then after lowering the pitch you would have [3, 5, 4, 0, 0]. So they all shifted a bin lower.  So from a convolutional perspective this doesn't really change the \"picture\" and as a result doesn't add much variety to the training set. But perhaps I overlooked something.",
      "votes": null
    },
    {
      "id": "268172",
      "postDate": "01/13/2018 16:23:08",
      "content": "<p>pitch <code>data augmentation</code> don't have any benefit in my solutions.</p>",
      "rawMarkdown": "pitch `data augmentation` don't have any benefit in my solutions.",
      "votes": null
    },
    {
      "id": "268181",
      "postDate": "01/13/2018 16:41:41",
      "content": "<p>i am thinking that you can apply it to only the unknown class only.</p>",
      "rawMarkdown": "i am thinking that you can apply it to only the unknown class only.",
      "votes": null
    },
    {
      "id": "268198",
      "postDate": "01/13/2018 18:01:29",
      "content": "<p>I used it for generating different silence samples from the background_noise files and it helped. But I found it didn't help if used as part of regular training augmentation.</p>",
      "rawMarkdown": "I used it for generating different silence samples from the background_noise files and it helped. But I found it didn't help if used as part of regular training augmentation.",
      "votes": null
    },
    {
      "id": "268218",
      "postDate": "01/13/2018 20:11:14",
      "content": "<p>I will try with silence</p>",
      "rawMarkdown": "I will try with silence",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 268113,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "01/13/2018 11:41:00",
      "content": "<p>For us, the pitch shift hurts LB and is also too slow, so we removed it from our data augmentation pipeline.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268142,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "01/13/2018 14:17:04",
      "content": "<p>I tried it and also didn't see improvements. When I thought about it a bit more, what would actually be happening is that you just shift the distribution over the frequency bins (assuming you use some kind of spectrogram feature). </p>\n\n<p>So if you had for example 5 frequency bins with the following (imaginary) values [0, 3, 5, 4, 0] then after lowering the pitch you would have [3, 5, 4, 0, 0]. So they all shifted a bin lower.  So from a convolutional perspective this doesn't really change the \"picture\" and as a result doesn't add much variety to the training set. But perhaps I overlooked something.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268172,
      "author_name": "feitengli",
      "author_url": "",
      "post_date": "01/13/2018 16:23:08",
      "content": "<p>pitch <code>data augmentation</code> don't have any benefit in my solutions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268181,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/13/2018 16:41:41",
      "content": "<p>i am thinking that you can apply it to only the unknown class only.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268198,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "01/13/2018 18:01:29",
      "content": "<p>I used it for generating different silence samples from the background_noise files and it helped. But I found it didn't help if used as part of regular training augmentation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 268218,
          "author_name": "princerk",
          "author_url": "",
          "post_date": "01/13/2018 20:11:14",
          "content": "<p>I will try with silence</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "268057": "Does any one got success with pitch shift data augmentation? I am using librosa to add  pitch shift in my samples but it does not seem to help much I have tried with other data augmentation like mixing noise and time shift that helped me gain 2% LB. In particularly I am using the following code\n\n    def pitch_shift(samples, sr=sr):\n        return librosa.effects.pitch_shift(samples, sr=sr, n_steps=random.randint(1, 5))\n\nDo I have to change some parameters? \nPercentage of data in training with pitch shift?\n\nedit 1:\nGiven the limited time to do detailed experiments. I have not tried time stretching. Does it help? \n\n    librosa.effects.time_stretch(y, rate)",
    "268113": "For us, the pitch shift hurts LB and is also too slow, so we removed it from our data augmentation pipeline.",
    "268142": "I tried it and also didn't see improvements. When I thought about it a bit more, what would actually be happening is that you just shift the distribution over the frequency bins (assuming you use some kind of spectrogram feature). \n\nSo if you had for example 5 frequency bins with the following (imaginary) values [0, 3, 5, 4, 0] then after lowering the pitch you would have [3, 5, 4, 0, 0]. So they all shifted a bin lower.  So from a convolutional perspective this doesn't really change the \"picture\" and as a result doesn't add much variety to the training set. But perhaps I overlooked something.",
    "268172": "pitch `data augmentation` don't have any benefit in my solutions.",
    "268181": "i am thinking that you can apply it to only the unknown class only.",
    "268198": "I used it for generating different silence samples from the background_noise files and it helped. But I found it didn't help if used as part of regular training augmentation.",
    "268218": "I will try with silence"
  },
  "source": "meta"
}