{"cells":[{"metadata":{},"cell_type":"markdown","source":"I took this competition as an opportunity to experiment with Dask and distributed computing. Everything was made using Kaggle kernels.\n\nIn a previous notebook I created a preprocessing scheme using **PyDub, Dask and Zarr** to read each MP3 file, resample it to 32MHz as recommended by this competition hosts, and extract only one audio channel.\n\nThe recording is then converted to a sequence stored as a NumPy array. Each audio sequence has been split into subsequences using a non overlapping moving window of size 46^3, or approximately 3 seconds of audio. These sequences are left padded with zeroes. The final shape of the array is (n,46,46,46), however it’s also possible to slice and reshape it using the current solution or even create a different preprocessing scheme.\n\nFor example the first recording has 815616 steps and is splitted into an array of 9 rows.\n\nThe final result is stored as a compressed Zarr array composed of approximately 677 chunks totaling 42gb and uploaded as a Kaggle dataset. \n\nIf you find this useful, I can share the kernel where it’s possible to adapt or alter the preprocessing to your own requirements.\n","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"!conda install dask -y\nimport dask\nimport dask.array as da\nimport dask.bag as db\nimport dask.dataframe as dd\nfrom dask.distributed import Client\n\n!pip install pydub\nfrom pydub import AudioSegment\nfrom pydub.utils import mediainfo\n\n!pip install zarr\nfrom zarr import Zlib, BZ2, LZMA, Blosc\n\nimport librosa, os, time, gc, json\nimport numpy as np\nimport h5py\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport IPython.display as ipd\nfrom skimage.util.shape import view_as_windows, view_as_blocks\nimport tensorflow as tf\n\npd.set_option(\"display.max_columns\", 60)\npd.set_option(\"display.max_rows\", 120)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Set up Dask Client","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"temp_folder = '/home/dask_hdd'\nos.mkdir(temp_folder)\n\nclient = Client(memory_limit='4GB', local_directory=temp_folder)\nclient","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Read Zarr dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"audio = da.from_zarr('/kaggle/input/bird-train')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"audio","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**It's possible to retrieve any part of the array relatively fast.**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nrecording = audio[20000:20500].compute()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"recording","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Create index\n\n\nInside the dataset bird-train there is also a copy of the train dataframe with two columns \"len\" and \"path\" included.\n\nThe \"len\" column has the exact length of each audio sequence, and can be used to create an index. The length does not include left padding, so this has to be accounted for in the code. Also, some mp3 files could not be read, these were filtered out.\n\nThis way it's possible to create an index that maps a specific line of the train dataframe to a slice of the zarr array that contains that recording (audio_index). The inverse operation is also possible, map any array row to a specific entry in the train dataframe (train_index). ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_feather('/kaggle/input/bird-train/train.feather')\ntrain = train[train['len']>0]\n\nwindow_maxdim = np.prod(audio.shape[1:])\ntrain['shape'] = train['len'].map(lambda x: (x+window_maxdim-x%window_maxdim)/window_maxdim)\ntrain['slice'] = train['shape'].cumsum().astype(int)\ntrain['slice'] = list(zip(train['slice'].shift(1,fill_value=0),train['slice']))\n\naudio_index = train['slice'].map(lambda x: slice(*x))\ntrain_index = train['slice'].apply(lambda x: np.arange(*x)).explode()\ntrain_index = pd.Series(train_index.index, index=train_index.astype(int))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Retrieve recording from Zarr array","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[5000]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"audio_index.loc[5000]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"recording = audio[audio_index.loc[5000]].compute()\nprint(recording.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"recording = recording.reshape(1,-1)\nipd.Audio(recording, rate=32000)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Double check to make sure we got the right one!**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[5000].path","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"recording_original = AudioSegment.from_mp3(train.loc[5000].path).set_frame_rate(32000).set_channels(1)\nrecording_original","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"recording = recording.ravel()\nrecording_original = np.array(recording_original.get_array_of_samples(), dtype=np.int32)\nnp.all(np.equal(recording[-recording_original.shape[0]:], recording_original))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Retrieve recording information from array row","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"audio[99553]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_index.loc[99553]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[train_index.loc[99553]]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Get max value of entire array\n\nGo through an entire 158gb array in under 2 minutes with only 16gb of ram!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\naudio.max().compute()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Get exact position of max value**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\naudio.argmax().compute()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nnp.unravel_index(1280089650, audio.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"audio[13151, 11, 13, 40].compute()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Bonus: group by recording and apply function\n\nGet min and max value of each recording.\n\nBe careful when calculating statistics like mean because the array has been zero padded. This can be accounted for during calculation. The processing notebook can also be modified to include numpy masked arrays.\n\nThe process takes about 3 minutes, but there may be more efficient ways to do this with Dask.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ngroupby_minmax = [dask.delayed(lambda x: (x.min(), x.max()))(audio[sli]) for sli in audio_index]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ngroupby_minmax = db.compute(groupby_minmax)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"groupby_minmax","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}