{"cells":[{"metadata":{"_uuid":"165585eca38e527a68b51e1b21c82c88eb02e0a6"},"cell_type":"markdown","source":"# About\nThis section is to intoduce lot of data augmentation techniques inspried by sound analysis techniques.\nStating from the basic cropping, we would include some more advanced stuffs, such as wave synthesize and pitch shifting etc.  "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nimport pyarrow.parquet as pq\nimport matplotlib.pyplot as plt\n\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"34c45601a89e55a320350768f4e044ce3cc36bfe"},"cell_type":"markdown","source":" ## Augmentors \nThese augmentors are inspired by https://arxiv.org/pdf/1711.10282.pdf , https://arxiv.org/pdf/1604.07160.pdf and http://karol.piczak.com/papers/Piczak2015-ESC-ConvNet.pdf. \n\nAlso, some of the techniques mentioned in https://arxiv.org/pdf/1608.04363.pdf can be useful, so I am looking to add these, too.\n\nFor the random cropping, I set the default length to be 50k.\nIn 2D CNN(images), in many applications,  the input size is $(224, 256)^2$,  which makes 45~ 65k data points.\n\nAlso, it is mentioned that 1s~2s of sound wave is large enough in https://arxiv.org/pdf/1711.10282.pdf. Considering the sampling rate of CD is 44.1 kHZ, that is equivalent to 44.1 ~ 88.2k data points. "},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"def random_crop(data,meta_data = None,length = 50000,n_samples = 100,targets = None,mode = 'uniform',sigma = 400000):\n    if meta_data is None:\n        meta_data = pd.read_csv('../input/metadata_train.csv')\n    if length is None:\n        length = np.random.randint(1,len(data))\n    if mode != 'uniform':\n        start = np.random.randn() * sigma + len(data) // 2\n    else:\n        start = np.random.randint(1,len(data))\n    if isinstance(targets,int):\n        meta_data = meta_data[meta_data['target'] == targets]\n    indices = range(len(meta_data))\n    #selected_data = data[:,indices]\n    if n_samples > meta_data.shape[0]:\n        n_samples = meta_data.shape[0]\n    inds = np.random.choice(indices,n_samples,replace = False)\n    targets = meta_data.iloc[inds].target.values\n    pair = (data[start:start + length,inds],targets)\n    return  pair\n     ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"102965fcef80d5a163c7f546e6bc6a3a103bd9cf"},"cell_type":"code","source":"# assumes random crop is done\ndef synthesize_waves(pair_a,pair_b,alpha = .5):\n    assert np.all(pair_a[1] == pair_b[1]), 'class does not match'\n    assert len(pair_a[0]) == len(pair_b[0]), 'wave length does not match'\n    np.random.shuffle(pair_b[0].T)\n    return alpha * pair_a[0] + (1 - alpha) * pair_b[0]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"12444d86e06487836552100811c5e6b1c71be464"},"cell_type":"code","source":"#test and show some results here\ndata  = pq.read_pandas('../input/train.parquet', columns=[str(i) for i in range(510)]).to_pandas().values\nmeta_data = pd.read_csv('../input/metadata_train.csv')[:510]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f1a72ccb1244da7b9098ef033e3e042d3b536ba","trusted":true},"cell_type":"code","source":"pair_a = random_crop(data,meta_data) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e23a4cfac67b52115bd5855d6be3f374911e4274"},"cell_type":"code","source":"plt.plot(pair_a[0][:,3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3a215f6d4ca61de36ff2786e9a8bae8642848266"},"cell_type":"code","source":"pair_b = random_crop(data,meta_data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"42977dd6ff85dd8a91987e73b4105975e771d511"},"cell_type":"code","source":"plt.plot(pair_b[0][:,3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"eb0373c23e209e176c0f077306b847e30bbf20d3"},"cell_type":"code","source":"#synthesize the waves\npair_a = random_crop(data,meta_data,targets = 0)\npair_b = random_crop(data,meta_data,targets = 0)\nsynthesized = synthesize_waves(pair_a,pair_b)\nplt.plot(synthesized[:,3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d62bbe556447cd2e8ebab13c99f7ac096ad38eed"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fd9db692feed034b622e4c6d6651444e95e05bb1"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}