{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"},{"sourceId":15853,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":2739,"modelId":319}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Audio Embeddings with Auditus\n\n\n![](https://www.kaggle.com/competitions/91844/images/header)\n\n[auditus](https://github.com/CarloLepelaars/auditus) is a new library for generating simple audio embeddings. In this notebook we will generate embeddings for the full training dataset using [Google's Bird Vocalization Classifier model](https://www.kaggle.com/models/google/bird-vocalization-classifier).","metadata":{}},{"cell_type":"markdown","source":"## Preparation","metadata":{}},{"cell_type":"code","source":"!pip install -Uqq auditus","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:55:53.865240Z","iopub.execute_input":"2025-03-26T15:55:53.865585Z","iopub.status.idle":"2025-03-26T15:55:57.450817Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom os.path import relpath\nfrom fastcore.all import *\nfrom fastprogress.fastprogress import progress_bar\nfrom fasttransform import Pipeline\n\nfrom auditus.core import AudioArray\nfrom auditus.transform import AudioLoader, Resampling, TFAudioEmbedding","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:00.648005Z","iopub.execute_input":"2025-03-26T15:56:00.648338Z","iopub.status.idle":"2025-03-26T15:56:07.639851Z","shell.execute_reply.started":"2025-03-26T15:56:00.648308Z","shell.execute_reply":"2025-03-26T15:56:07.638910Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Kaggle caps the Notebook output at 500 files so we will add `train_audio` embeddings as a column to the `train.csv` file.","metadata":{}},{"cell_type":"code","source":"BASE_PATH = \"/kaggle/input/birdclef-2025/\"\ntrain_soundscape_paths = np.random.choice(globtastic(f\"{BASE_PATH}train_soundscapes\", file_glob=\"*.ogg\"), size=250)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:07.641050Z","iopub.execute_input":"2025-03-26T15:56:07.641563Z","iopub.status.idle":"2025-03-26T15:56:11.400865Z","shell.execute_reply.started":"2025-03-26T15:56:07.641539Z","shell.execute_reply":"2025-03-26T15:56:11.400186Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_soundscape_paths[:3]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:11.402299Z","iopub.execute_input":"2025-03-26T15:56:11.402529Z","iopub.status.idle":"2025-03-26T15:56:11.408604Z","shell.execute_reply.started":"2025-03-26T15:56:11.402509Z","shell.execute_reply":"2025-03-26T15:56:11.407656Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df = pd.read_csv(f\"{BASE_PATH}train.csv\")\ndf.head(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:11.409463Z","iopub.execute_input":"2025-03-26T15:56:11.409730Z","iopub.status.idle":"2025-03-26T15:56:11.536223Z","shell.execute_reply.started":"2025-03-26T15:56:11.409709Z","shell.execute_reply":"2025-03-26T15:56:11.535306Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_audio_paths = [f\"{BASE_PATH}train_audio/{name}\" for name in df['filename']]\nprint(len(train_audio_paths))\ntrain_audio_paths[:3]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:11.537062Z","iopub.execute_input":"2025-03-26T15:56:11.537387Z","iopub.status.idle":"2025-03-26T15:56:11.551388Z","shell.execute_reply.started":"2025-03-26T15:56:11.537362Z","shell.execute_reply":"2025-03-26T15:56:11.550186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Single instance","metadata":{}},{"cell_type":"markdown","source":"Let's look at a single example first.\n\nThe Bird vocalization model we use only works on $5$ seconds of data. For simplicity we truncate each file to the first $5$ seconds. The audio files have a sample rate of $32000$ so for $5$ seconds this equates to a length of $160000$.","metadata":{}},{"cell_type":"code","source":"class Truncate(Transform):\n    \"\"\" Get first 5 seconds for sample rate of 32000 and add padding if necessary. \"\"\"\n    def encodes(self, x:AudioArray): return AudioArray(np.pad(x.a, (0, max(0, 160000-len(x.a))))[:160000], x.sr)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:11.552282Z","iopub.execute_input":"2025-03-26T15:56:11.552542Z","iopub.status.idle":"2025-03-26T15:56:11.566954Z","shell.execute_reply.started":"2025-03-26T15:56:11.552520Z","shell.execute_reply":"2025-03-26T15:56:11.566208Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_path = np.random.choice(train_audio_paths)\nprint(f\"Test file: '{test_path}'\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:17.488361Z","iopub.execute_input":"2025-03-26T15:56:17.488695Z","iopub.status.idle":"2025-03-26T15:56:17.504303Z","shell.execute_reply.started":"2025-03-26T15:56:17.488669Z","shell.execute_reply":"2025-03-26T15:56:17.503392Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Full file:","metadata":{}},{"cell_type":"code","source":"full_audio = AudioLoader(sr=32000)(test_path)\nfull_audio.audio()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:18.002031Z","iopub.execute_input":"2025-03-26T15:56:18.002374Z","iopub.status.idle":"2025-03-26T15:56:18.394778Z","shell.execute_reply.started":"2025-03-26T15:56:18.002346Z","shell.execute_reply":"2025-03-26T15:56:18.393267Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After truncation:","metadata":{}},{"cell_type":"code","source":"truncated_audio = Truncate()(full_audio)\ntruncated_audio.audio()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:18.995206Z","iopub.execute_input":"2025-03-26T15:56:18.995551Z","iopub.status.idle":"2025-03-26T15:56:19.021951Z","shell.execute_reply.started":"2025-03-26T15:56:18.995519Z","shell.execute_reply":"2025-03-26T15:56:19.021079Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Pipeline","metadata":{}},{"cell_type":"markdown","source":"We start with a random audio file. The pipeline generates an embedding of length 1280. ","metadata":{}},{"cell_type":"code","source":"pipe = Pipeline([AudioLoader(sr=32000), Truncate(), TFAudioEmbedding('/kaggle/input/bird-vocalization-classifier/tensorflow2/bird-vocalization-classifier/8/')])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:56:26.207204Z","iopub.execute_input":"2025-03-26T15:56:26.207518Z","iopub.status.idle":"2025-03-26T15:56:30.109327Z","shell.execute_reply.started":"2025-03-26T15:56:26.207495Z","shell.execute_reply":"2025-03-26T15:56:30.108660Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"output = pipe(test_path)\nprint(f\"Embedding Length: '{output.shape}'\")\nprint(f\"First 5 elements of embedding: '{output[0][:5]}'\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-26T15:59:52.130170Z","iopub.execute_input":"2025-03-26T15:59:52.130526Z","iopub.status.idle":"2025-03-26T15:59:52.287448Z","shell.execute_reply.started":"2025-03-26T15:59:52.130500Z","shell.execute_reply":"2025-03-26T15:59:52.286518Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Full dataset","metadata":{}},{"cell_type":"markdown","source":"We are now ready to generate embeddings for the full training dataset. Using the GPU this takes around 10 minutes for the full training set.","metadata":{}},{"cell_type":"code","source":"def get_emb(path): return pipe(path).squeeze(0).tolist()\n\nwith ThreadPoolExecutor() as ex:\n    train_embs = list(progress_bar(ex.map(get_emb, train_audio_paths), total=len(train_audio_paths)))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df[\"emb\"] = train_embs","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.to_csv(\"train_with_emb.csv\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The train file with embeddings is stored as a [Kaggle Dataset for easy exploration](https://www.kaggle.com/datasets/carlolepelaars/birdclef-2025-perch-embeddings). In [the next notebook](https://www.kaggle.com/code/carlolepelaars/birdclef-2025-embedding-visualization) we visualize these embeddings in 2 dimensions with [UMAP](https://umap-learn.readthedocs.io/en/latest) to get a good overview of the data and find patterns between the classes.","metadata":{}},{"cell_type":"markdown","source":"**That's it! Hope this helps you to get started with the competition!**\n\n**If you like this Kaggle Notebook, consider giving an upvote and leaving a comment. Your feedback is very welcome! I will try to implement your suggestions.**","metadata":{}}]}