{"cells":[{"metadata":{},"cell_type":"markdown","source":"## <center>Augmentation methods for audio\nIntuitively, lack of data is one of the common issue in actual data science problem. Data augmentation helps to generate synthetic data from existing data set such that generalisation capability of model can be improved.\n  <p> <b>Data augmentation definition :</b>\n<ul>\n  <li>Data augmentation is the process by which we create new synthetic training samples by adding small perturbations on our initial training set.</li>\n  <li>The objective is to make our model invariant to those perturbations and enhace its ability to generalize.</li>\n  <li>In images data augmention can be performed by shifting the image, zooming, rotating ...</li>\n    <li>In our case we will add noise, stretch and roll, pitch shift ...</li>\n</ul>\n","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport librosa\nimport matplotlib.pyplot as plt\nimport os\nimport cv2\nimport IPython.display as ipd\nfrom IPython.display import Audio, IFrame, display\nimport plotly.graph_objects as go\nimport librosa\nimport librosa.display\nplt.style.use(\"ggplot\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train = pd.read_csv(\"../input/birdsong-recognition/train.csv\")\nspecies=train.species.value_counts()\nfig = go.Figure(data=[\n    go.Bar(y=species.values, x=species.index,marker_color='deeppink')\n])\n\nfig.update_layout(title='Distribution of Bird Species')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path='../input/birdsong-resampled-train-audio-04/wooscj2/XC67042.wav'\nx , sr = librosa.load(file_path)\nlibrosa.display.waveplot(x, sr=sr)\nAudio(x, rate=sr)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1. Noise Injection\nIt simply add some random value into data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def noise(data, noise_factor):\n    noise = np.random.randn(len(data))\n    augmented_data = data + noise_factor * noise\n    # Cast back to same data type\n    augmented_data = augmented_data.astype(type(data[0]))\n    return augmented_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n=noise(x,0.01)\nlibrosa.display.waveplot(n, sr=sr)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 2. Time shifting\nslightly shift the starting point of the audio, then pad it to original length.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def shifting_time(data, sampling_rate, shift_max, shift_direction):\n    shift = np.random.randint(sampling_rate * shift_max)\n    if shift_direction == 'right':\n        shift = -shift\n    elif self.shift_direction == 'both':\n        direction = np.random.randint(0, 2)\n        if direction == 1:\n            shift = -shift\n    augmented_data = np.roll(data, shift)\n    # Set to silence for heading/ tailing\n    if shift > 0:\n        augmented_data[:shift] = 0\n    else:\n        augmented_data[shift:] = 0\n    return augmented_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"s=shifting_time(x,sr,1,'right')\nlibrosa.display.waveplot(s, sr=sr)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 3. Speed tuning\nslightly change the speed of the audio, then pad or slice it.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def speed(data, speed_factor):\n    return librosa.effects.time_stretch(data, speed_factor)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"v=speed(x,2)\nlibrosa.display.waveplot(v, sr=sr)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 4. Changing Pitch","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def pitch(data, sampling_rate, pitch_factor):\n    return librosa.effects.pitch_shift(data, sampling_rate, pitch_factor)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"p=pitch(x,sr,2)\nlibrosa.display.waveplot(p, sr)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Take Away\n<ul>\n  <li>Data augmentation cannot replace real training data. It just help to generate synthetic data to make the model better.\n</li>\n  <li>Do not blindly generate synthetic data. You have to understand your data pattern and selecting a appropriate way to increase training data volume.</li>\n</ul>\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Please do upvote if you find it useful 👍🏼","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}