{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [BirdCLEF23](https://www.kaggle.com/competitions/birdclef-2023) Audio to Numpy\n\n## Create NumPy arrays for the sound-files\n\n[Adapted from here](https://www.kaggle.com/code/kaerunantoka/birdclef2022-audio-to-numpy-1-4).\n\n- This notebook turned out to be a dead-end, I've decided to just load .ogg files on the fly during training instead.\n\n- I'm mainly putting this up to save others some time.  The thought was that by loading numpy arrays during training might save some training time.  But from my little experiments, I don't think it's going to help.  Using the BirdCLEF-2023 dataset I got the following results:\n\n- Loading the full dataset from .ogg files one at a time, and converting each to NumPy arrays: **2:29**\n\n- Loading the full dataset directly from .npy arrays one at a time: **2:47**.\n\n- The datasets end up crazy large and you'll run into Kaggle disk space limits, but it works OK on a local machine. On my Dell G15 the full dataset 5GB unpacks in 3:36 seconds.  \n\n- **Output:** The 2023 BirdCLEF sound files in NumPy .npy format, with the same file structure as the original dataset.  **Warning** This new dataset will be **151.0 GB**.\n\n- **Usage:** You should just need to modify the two filepaths `data_folder` and `out_folder`, and un-comment the last line. (I've commented it out so I can commit without crashing the notebook due to memory limits).\n\n- Curious to see if I've made any errors in this analysis.  Feel free to comment.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-20T05:01:26.897960Z","iopub.execute_input":"2023-03-20T05:01:26.899120Z","iopub.status.idle":"2023-03-20T05:01:26.908618Z","shell.execute_reply.started":"2023-03-20T05:01:26.899074Z","shell.execute_reply":"2023-03-20T05:01:26.906979Z"}}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport soundfile as sf\nfrom joblib import Parallel, delayed\nfrom tqdm import tqdm\nfrom pathlib import Path\nfrom os import sep\nimport glob\nimport gc","metadata":{"execution":{"iopub.status.busy":"2023-03-24T02:09:03.946227Z","iopub.execute_input":"2023-03-24T02:09:03.946686Z","iopub.status.idle":"2023-03-24T02:09:03.952906Z","shell.execute_reply.started":"2023-03-24T02:09:03.946650Z","shell.execute_reply":"2023-03-24T02:09:03.951421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NUM_WORKERS = 4\nSR = 32000 # Sampling rate of all the source files\nUSE_SEC = 300  # Clips the files to this number of seconds.","metadata":{"execution":{"iopub.status.busy":"2023-03-24T02:09:03.955552Z","iopub.execute_input":"2023-03-24T02:09:03.955907Z","iopub.status.idle":"2023-03-24T02:09:03.979650Z","shell.execute_reply.started":"2023-03-24T02:09:03.955874Z","shell.execute_reply":"2023-03-24T02:09:03.978479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Setup filepaths and create output folders\n\n- Write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n- You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{}},{"cell_type":"code","source":"data_folder = Path('/kaggle/input')  # modify to suit\nout_folder = Path('/kaggle/temp') # will be lost outside current session\n\nnp_out = out_folder / 'np_audio_23'\ntrain_23 = data_folder / 'birdclef-2023' / 'train_audio'\npaths = {p:str(p.name) for p in Path(train_23).rglob('*.ogg')}\nclasses = os.listdir(train_23)\n\nos.makedirs(np_out, exist_ok=True)\nfor bird_type in tqdm(classes):\n    os.makedirs(np_out / bird_type, exist_ok=True)\n\nprint(f'There are a total of {len(classes)} bird class folders')\nprint(f'There are a total of {len(paths)} bird sound files')","metadata":{"execution":{"iopub.status.busy":"2023-03-24T02:09:03.981312Z","iopub.execute_input":"2023-03-24T02:09:03.981623Z","iopub.status.idle":"2023-03-24T02:09:04.381304Z","shell.execute_reply.started":"2023-03-24T02:09:03.981593Z","shell.execute_reply":"2023-03-24T02:09:04.380126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Helper functions","metadata":{}},{"cell_type":"code","source":"def Audio_to_Array(path):\n    y, sr = sf.read(path, always_2d=True)\n    y = np.mean(y, 1) # For any sterio (X, 2) arrays\n    if len(y) > SR:\n        y = y[SR:-SR]\n\n    if len(y) > SR * USE_SEC:\n        y = y[:SR * USE_SEC]\n    return y\n\n\ndef save_(path):\n    save_folder = str(path.parent.name)\n    save_name = str(path.stem) + '.npy'\n    save_path = np_out / save_folder / save_name\n    sf =  Audio_to_Array(path)\n    np.save(save_path, sf)\n    sf = None # to release memory\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T02:09:04.385125Z","iopub.execute_input":"2023-03-24T02:09:04.385884Z","iopub.status.idle":"2023-03-24T02:09:04.394242Z","shell.execute_reply.started":"2023-03-24T02:09:04.385830Z","shell.execute_reply":"2023-03-24T02:09:04.393333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Open all the ogg files and save them as numpy arrays  Uncomment the line below.","metadata":{}},{"cell_type":"code","source":"#_ = Parallel(n_jobs=NUM_WORKERS)(delayed(save_)(path) for path in tqdm(paths))","metadata":{"execution":{"iopub.status.busy":"2023-03-24T02:09:04.397410Z","iopub.execute_input":"2023-03-24T02:09:04.398399Z","iopub.status.idle":"2023-03-24T02:09:04.418589Z","shell.execute_reply.started":"2023-03-24T02:09:04.398339Z","shell.execute_reply":"2023-03-24T02:09:04.417366Z"},"trusted":true},"execution_count":null,"outputs":[]}]}