{
  "id": 170913,
  "title": "[Imp] Augmentation methods for audio",
  "url": "/competitions/birdsong-recognition/discussion/170913",
  "author_name": "",
  "post_date": "2020-07-29T14:03:01.623779400Z",
  "votes": 22,
  "comment_count": 9,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/jainarindam/audio-augmentation-methods\">code</a></p>\n\n<h2>Data augmentation definition :</h2>\n\n<p>Data augmentation is the process by which we create new synthetic training samples by adding small perturbations on our initial training set.\nThe objective is to make our model invariant to those perturbations and enhace its ability to generalize.\nIn images data augmention can be performed by shifting the image, zooming, rotating ...\nIn our case we will add noise, stretch and roll, pitch shift ...</p>\n\n<h2>1. Adding Noise</h2>\n\n<p>It simply add some random value into data by using numpy.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4df0b043457402ad7318d109cc5a3e80%2Fa4.png?generation=1596031716930344&amp;alt=media\" alt=\"\"></p>\n\n<h2>2. Shifting Time</h2>\n\n<p>It just shift audio to left/right with a random second.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F81a5bb5f944d4484074e37df47c1ef99%2Fa1.png?generation=1596031680234729&amp;alt=media\" alt=\"\"></p>\n\n<h2>3. Change Pitch</h2>\n\n<p>It Change pitch randomly\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fe37d3982575221941d893cbbf655b32c%2Fa2.png?generation=1596031641223863&amp;alt=media\" alt=\"\"></p>\n\n<h2>4. Change Speed</h2>\n\n<p>It stretches times series by a fixed rate.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fc0b6dd5bef633c0bd747eda7ed57e228%2Fa3.png?generation=1596031609456536&amp;alt=media\" alt=\"\"></p>\n\n<h3>If you have more ideas please let me know in discussion I will add them</h3>\n\n<p>Using these techniques we can create more data and randomness so it can generalise the results </p>\n\n<p><a href=\"https://www.kaggle.com/jainarindam/audio-augmentation-methods\">Check the code here</a></p>",
  "messages": [
    {
      "id": "950630",
      "postDate": "07/29/2020 14:03:01",
      "content": "<p><a href=\"https://www.kaggle.com/jainarindam/audio-augmentation-methods\">code</a></p>\n\n<h2>Data augmentation definition :</h2>\n\n<p>Data augmentation is the process by which we create new synthetic training samples by adding small perturbations on our initial training set.\nThe objective is to make our model invariant to those perturbations and enhace its ability to generalize.\nIn images data augmention can be performed by shifting the image, zooming, rotating ...\nIn our case we will add noise, stretch and roll, pitch shift ...</p>\n\n<h2>1. Adding Noise</h2>\n\n<p>It simply add some random value into data by using numpy.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4df0b043457402ad7318d109cc5a3e80%2Fa4.png?generation=1596031716930344&amp;alt=media\" alt=\"\"></p>\n\n<h2>2. Shifting Time</h2>\n\n<p>It just shift audio to left/right with a random second.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F81a5bb5f944d4484074e37df47c1ef99%2Fa1.png?generation=1596031680234729&amp;alt=media\" alt=\"\"></p>\n\n<h2>3. Change Pitch</h2>\n\n<p>It Change pitch randomly\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fe37d3982575221941d893cbbf655b32c%2Fa2.png?generation=1596031641223863&amp;alt=media\" alt=\"\"></p>\n\n<h2>4. Change Speed</h2>\n\n<p>It stretches times series by a fixed rate.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fc0b6dd5bef633c0bd747eda7ed57e228%2Fa3.png?generation=1596031609456536&amp;alt=media\" alt=\"\"></p>\n\n<h3>If you have more ideas please let me know in discussion I will add them</h3>\n\n<p>Using these techniques we can create more data and randomness so it can generalise the results </p>\n\n<p><a href=\"https://www.kaggle.com/jainarindam/audio-augmentation-methods\">Check the code here</a></p>",
      "rawMarkdown": "[code](https://www.kaggle.com/jainarindam/audio-augmentation-methods)\n\n## Data augmentation definition :\n\nData augmentation is the process by which we create new synthetic training samples by adding small perturbations on our initial training set.\nThe objective is to make our model invariant to those perturbations and enhace its ability to generalize.\nIn images data augmention can be performed by shifting the image, zooming, rotating ...\nIn our case we will add noise, stretch and roll, pitch shift ...\n\n## 1. Adding Noise\nIt simply add some random value into data by using numpy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4df0b043457402ad7318d109cc5a3e80%2Fa4.png?generation=1596031716930344&amp;alt=media)\n\n\n## 2. Shifting Time \nIt just shift audio to left/right with a random second.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F81a5bb5f944d4484074e37df47c1ef99%2Fa1.png?generation=1596031680234729&amp;alt=media)\n\n## 3. Change Pitch\nIt Change pitch randomly\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fe37d3982575221941d893cbbf655b32c%2Fa2.png?generation=1596031641223863&amp;alt=media)\n\n## 4. Change Speed\n It stretches times series by a fixed rate.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fc0b6dd5bef633c0bd747eda7ed57e228%2Fa3.png?generation=1596031609456536&amp;alt=media)\n\n### If you have more ideas please let me know in discussion I will add them\nUsing these techniques we can create more data and randomness so it can generalise the results \n\n [Check the code here](https://www.kaggle.com/jainarindam/audio-augmentation-methods)",
      "votes": null
    },
    {
      "id": "950644",
      "postDate": "07/29/2020 14:14:05",
      "content": "<p>I was thinking of using mirror and up-side down in Audio as we used to do in image augmentation. I am not sure if that will work. What are your thought on that?</p>",
      "rawMarkdown": "I was thinking of using mirror and up-side down in Audio as we used to do in image augmentation. I am not sure if that will work. What are your thought on that?",
      "votes": null
    },
    {
      "id": "965075",
      "postDate": "08/10/2020 11:19:44",
      "content": "<p>I am not sure but if I had to add another one to this it could be Volume Tuning.\nLet me know if I am right or not.</p>",
      "rawMarkdown": "I am not sure but if I had to add another one to this it could be Volume Tuning.\nLet me know if I am right or not.",
      "votes": null
    },
    {
      "id": "965501",
      "postDate": "08/10/2020 17:18:52",
      "content": "<p>Just like in Car :p \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F5b50bc633e703c447afbe19e27571910%2Fcarrrr.jpg?generation=1597079779424528&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Just like in Car :p \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F5b50bc633e703c447afbe19e27571910%2Fcarrrr.jpg?generation=1597079779424528&amp;alt=media)",
      "votes": null
    },
    {
      "id": "972420",
      "postDate": "08/16/2020 14:41:35",
      "content": "<p>normalizing bird sound ( to make predictions for faint sounds). This is discussed here: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/171133\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/171133</a><br>\nlooking forward to your implementation </p>",
      "rawMarkdown": "normalizing bird sound ( to make predictions for faint sounds). This is discussed here: https://www.kaggle.com/c/birdsong-recognition/discussion/171133\nlooking forward to your implementation",
      "votes": null
    },
    {
      "id": "988218",
      "postDate": "08/27/2020 22:11:35",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jainarindam\" target=\"_blank\">@jainarindam</a> I have created a notebook with all my ideas and implemented transforms on top of albumentations library making it easy to use <br>\njust do albumenatations.compose<br>\n<a href=\"https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio\" target=\"_blank\">https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio</a><br>\nI hope this helps</p>",
      "rawMarkdown": "Hey @jainarindam I have created a notebook with all my ideas and implemented transforms on top of albumentations library making it easy to use \njust do albumenatations.compose\nhttps://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio\nI hope this helps",
      "votes": null
    },
    {
      "id": "990938",
      "postDate": "08/30/2020 01:26:12",
      "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> awesome </p>",
      "rawMarkdown": "tanulsingh077 awesome",
      "votes": null
    },
    {
      "id": "1001356",
      "postDate": "09/07/2020 09:19:08",
      "content": "<p>Thanks for sharing Arindam. In your experience what has been the appropriate noise level to add for training ? 0.01 or something higher</p>",
      "rawMarkdown": "Thanks for sharing Arindam. In your experience what has been the appropriate noise level to add for training ? 0.01 or something higher",
      "votes": null
    },
    {
      "id": "1001735",
      "postDate": "09/07/2020 14:23:26",
      "content": "<p>Well I did it with higher that &gt;0.1 got no luck. probably less should work</p>",
      "rawMarkdown": "Well I did it with higher that >0.1 got no luck. probably less should work",
      "votes": null
    },
    {
      "id": "1001802",
      "postDate": "09/07/2020 15:27:09",
      "content": "<p>I tried 0.03 and it didn't work for me. </p>",
      "rawMarkdown": "I tried 0.03 and it didn't work for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 972420,
      "author_name": "ddbishnoi",
      "author_url": "",
      "post_date": "08/16/2020 14:41:35",
      "content": "<p>normalizing bird sound ( to make predictions for faint sounds). This is discussed here: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/171133\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/171133</a><br>\nlooking forward to your implementation </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 988218,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "08/27/2020 22:11:35",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jainarindam\" target=\"_blank\">@jainarindam</a> I have created a notebook with all my ideas and implemented transforms on top of albumentations library making it easy to use <br>\njust do albumenatations.compose<br>\n<a href=\"https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio\" target=\"_blank\">https://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio</a><br>\nI hope this helps</p>",
      "votes": null,
      "replies": [
        {
          "id": 990938,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "08/30/2020 01:26:12",
          "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> awesome </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1001356,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "09/07/2020 09:19:08",
      "content": "<p>Thanks for sharing Arindam. In your experience what has been the appropriate noise level to add for training ? 0.01 or something higher</p>",
      "votes": null,
      "replies": [
        {
          "id": 1001735,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "09/07/2020 14:23:26",
          "content": "<p>Well I did it with higher that &gt;0.1 got no luck. probably less should work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001802,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "09/07/2020 15:27:09",
          "content": "<p>I tried 0.03 and it didn't work for me. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 950644,
      "author_name": "jainarindam",
      "author_url": "",
      "post_date": "07/29/2020 14:14:05",
      "content": "<p>I was thinking of using mirror and up-side down in Audio as we used to do in image augmentation. I am not sure if that will work. What are your thought on that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 965075,
          "author_name": "palaksood97",
          "author_url": "",
          "post_date": "08/10/2020 11:19:44",
          "content": "<p>I am not sure but if I had to add another one to this it could be Volume Tuning.\nLet me know if I am right or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965501,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "08/10/2020 17:18:52",
          "content": "<p>Just like in Car :p \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F5b50bc633e703c447afbe19e27571910%2Fcarrrr.jpg?generation=1597079779424528&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "950630": "[code](https://www.kaggle.com/jainarindam/audio-augmentation-methods)\n\n## Data augmentation definition :\n\nData augmentation is the process by which we create new synthetic training samples by adding small perturbations on our initial training set.\nThe objective is to make our model invariant to those perturbations and enhace its ability to generalize.\nIn images data augmention can be performed by shifting the image, zooming, rotating ...\nIn our case we will add noise, stretch and roll, pitch shift ...\n\n## 1. Adding Noise\nIt simply add some random value into data by using numpy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4df0b043457402ad7318d109cc5a3e80%2Fa4.png?generation=1596031716930344&amp;alt=media)\n\n\n## 2. Shifting Time \nIt just shift audio to left/right with a random second.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F81a5bb5f944d4484074e37df47c1ef99%2Fa1.png?generation=1596031680234729&amp;alt=media)\n\n## 3. Change Pitch\nIt Change pitch randomly\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fe37d3982575221941d893cbbf655b32c%2Fa2.png?generation=1596031641223863&amp;alt=media)\n\n## 4. Change Speed\n It stretches times series by a fixed rate.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fc0b6dd5bef633c0bd747eda7ed57e228%2Fa3.png?generation=1596031609456536&amp;alt=media)\n\n### If you have more ideas please let me know in discussion I will add them\nUsing these techniques we can create more data and randomness so it can generalise the results \n\n [Check the code here](https://www.kaggle.com/jainarindam/audio-augmentation-methods)",
    "950644": "I was thinking of using mirror and up-side down in Audio as we used to do in image augmentation. I am not sure if that will work. What are your thought on that?",
    "965075": "I am not sure but if I had to add another one to this it could be Volume Tuning.\nLet me know if I am right or not.",
    "965501": "Just like in Car :p \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F5b50bc633e703c447afbe19e27571910%2Fcarrrr.jpg?generation=1597079779424528&amp;alt=media)",
    "972420": "normalizing bird sound ( to make predictions for faint sounds). This is discussed here: https://www.kaggle.com/c/birdsong-recognition/discussion/171133\nlooking forward to your implementation",
    "988218": "Hey @jainarindam I have created a notebook with all my ideas and implemented transforms on top of albumentations library making it easy to use \njust do albumenatations.compose\nhttps://www.kaggle.com/tanulsingh077/audio-albumentations-transform-your-audio\nI hope this helps",
    "990938": "tanulsingh077 awesome",
    "1001356": "Thanks for sharing Arindam. In your experience what has been the appropriate noise level to add for training ? 0.01 or something higher",
    "1001735": "Well I did it with higher that >0.1 got no luck. probably less should work",
    "1001802": "I tried 0.03 and it didn't work for me."
  },
  "source": "meta"
}