{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BirdCLEF+ 2025: Simple Submission","metadata":{}},{"cell_type":"markdown","source":"![](https://www.kaggle.com/competitions/91844/images/header)\n\nThis notebook shows a simple way to setup an inference pipeline for the [BirdCLEF+ 2025 competition](https://www.kaggle.com/competitions/birdclef-2025). \n\nCredits to [Stefan Kahl](https://www.kaggle.com/stefankahl) who set up [one of the first sample submission notebooks](https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission). \n\nI simplified the process and optimized the loading/chunking process with [numpy](https://numpy.org/) and [soundfile](https://github.com/bastibe/python-soundfile). It can be helpful to first think about the problem yourself and then check out both approaches.","metadata":{}},{"cell_type":"markdown","source":"## Dependencies","metadata":{}},{"cell_type":"markdown","source":"All these dependencies are included in the standard Kaggle Notebooks environment.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport soundfile as sf\n# Extension to pathlib.Path to simplify parsing directories\nfrom fastcore.xtras import Path","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-18T10:20:27.341761Z","iopub.execute_input":"2025-03-18T10:20:27.342094Z","iopub.status.idle":"2025-03-18T10:20:27.346508Z","shell.execute_reply.started":"2025-03-18T10:20:27.342064Z","shell.execute_reply":"2025-03-18T10:20:27.345373Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Get Paths and Labels","metadata":{}},{"cell_type":"markdown","source":"From the [competition description](https://www.kaggle.com/competitions/birdclef-2025/data) we know that the data is resampled to 32kHz, so we use a sample rate of `32_000`. A submission requires us to submit in 5 second chunks, so we chunk on `32000 * 5 = 160,000` samples.","metadata":{}},{"cell_type":"code","source":"BASE_PATH = \"/kaggle/input/birdclef-2025/\"\nTAXONOMY_PATH = f\"{BASE_PATH}taxonomy.csv\"\nTEST_SCAPES_PATH = f\"{BASE_PATH}test_soundscapes/\"\nSR, CHUNK_SEC = 32_000, 5\nFIVE_SEC_SR = SR * CHUNK_SEC","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-18T10:20:28.827709Z","iopub.execute_input":"2025-03-18T10:20:28.828021Z","iopub.status.idle":"2025-03-18T10:20:28.832557Z","shell.execute_reply.started":"2025-03-18T10:20:28.827996Z","shell.execute_reply":"2025-03-18T10:20:28.831425Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We retrieve the labels from the taxonomy file.","metadata":{}},{"cell_type":"code","source":"t = pd.read_csv(TAXONOMY_PATH)\nclass_labels = list(t['primary_label'])\nclass_labels[:5], class_labels[-5:]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-18T10:20:30.426719Z","iopub.execute_input":"2025-03-18T10:20:30.427018Z","iopub.status.idle":"2025-03-18T10:20:30.447706Z","shell.execute_reply.started":"2025-03-18T10:20:30.426994Z","shell.execute_reply":"2025-03-18T10:20:30.446806Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"`TEST_SCAPES_PATH` is populated during submission of the notebook.","metadata":{}},{"cell_type":"code","source":"# Populated during submission of notebook\nscape_paths = Path(TEST_SCAPES_PATH).ls(file_exts=\".ogg\")\nscape_paths[:10]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-18T10:21:47.981566Z","iopub.execute_input":"2025-03-18T10:21:47.981881Z","iopub.status.idle":"2025-03-18T10:21:47.988753Z","shell.execute_reply.started":"2025-03-18T10:21:47.981858Z","shell.execute_reply":"2025-03-18T10:21:47.987777Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Helper functions","metadata":{}},{"cell_type":"markdown","source":"Loading of audio is done efficiently using `soundfile`. `np.array_split` is an efficient way to split data into chunks.","metadata":{}},{"cell_type":"code","source":"def load_audio(path) -> np.array:\n    with sf.SoundFile(path) as f: audio = f.read()\n    return audio\n\ndef get_chunks(path) -> list[np.array]:\n    \"\"\" Create 5 second chunks (1D arrays) from audio file. \"\"\"\n    audio = load_audio(path)\n    return np.array_split(audio, np.ceil(audio.shape[0] / FIVE_SEC_SR))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T20:35:24.968064Z","iopub.execute_input":"2025-03-14T20:35:24.968507Z","iopub.status.idle":"2025-03-14T20:35:24.976840Z","shell.execute_reply.started":"2025-03-14T20:35:24.968464Z","shell.execute_reply":"2025-03-14T20:35:24.975582Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Inference","metadata":{}},{"cell_type":"markdown","source":"We make a prediction for each chunk within each soundscape.","metadata":{}},{"cell_type":"code","source":"cols = [\"row_id\"] + class_labels\nrows = []\nfor path in scape_paths:\n    for i, chunk in enumerate(get_chunks(path), start=1):\n        row_id = f\"{Path(path).stem}_{i * CHUNK_SEC}\"\n        # Place your inference function here (Random predictions as placeholder)\n        #########################################\n        pred = np.random.uniform(low=0.0, high=0.72, size=len(class_labels))\n        #########################################\n        rows.append([row_id] + pred.tolist())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T20:35:24.978154Z","iopub.execute_input":"2025-03-14T20:35:24.978711Z","iopub.status.idle":"2025-03-14T20:35:24.995038Z","shell.execute_reply.started":"2025-03-14T20:35:24.978667Z","shell.execute_reply":"2025-03-14T20:35:24.993766Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We make sure that the final file contains `row_id` and the `206` class labels. Be careful if you have shuffled the order of your class labels in training.","metadata":{}},{"cell_type":"code","source":"preds = pd.DataFrame(rows, columns=cols)\npreds.head(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T20:35:24.996383Z","iopub.execute_input":"2025-03-14T20:35:24.996752Z","iopub.status.idle":"2025-03-14T20:35:25.039587Z","shell.execute_reply.started":"2025-03-14T20:35:24.996727Z","shell.execute_reply":"2025-03-14T20:35:25.038393Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Submission","metadata":{}},{"cell_type":"code","source":"preds.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T20:35:25.040452Z","iopub.execute_input":"2025-03-14T20:35:25.040819Z","iopub.status.idle":"2025-03-14T20:35:25.051797Z","shell.execute_reply.started":"2025-03-14T20:35:25.040785Z","shell.execute_reply":"2025-03-14T20:35:25.050313Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**That's it! Hope this helps you to get started with the competition!**\n\n**If you like this Kaggle kernel, consider giving an upvote and leaving a comment. Your feedback is very welcome! I will try to implement your suggestions in this kernel.**","metadata":{}}]}