{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Working with BirdSet and Huggingface Datasets \n\n**BirdSet** offers a streamlined approach to accessing bird sound data. Here's what you need to know:\n\n- Formatted test and training datasets: Past BirdCLEF test datasets prepared in a 5-second evaluation format, with an alternative multi-class variant available.\n- Training subsets are matched to the test datasets: training data from XC is aligned with species in the test datasets.\n- We also offer big training datasets without test datasets. XCM covers all bird species from XC that appear in our benchmark test datasets and XCL is a complete snapshot from XC. \n- You could also only download the training or the test data if you do not need both (name_xc or name_scape) \n- We also provide detected events and clusters for each recording. \n\n- HF provides useful tutorials for working with audio datasets: https://huggingface.co/docs/datasets/en/audio_load","metadata":{}},{"cell_type":"markdown","source":"### Formats\n\n**train**: \n- Focal recordings from XC with quality ratings A, B, C and excluding all recordings that are CC-ND. \n- Each dataset is tailored for specific target species identified in the corresponding test soundscape files. \n- XCM and XCL provide bigger training datasets for pretraining. \n\n**test_5s**:\n- Task: Multilabel (\"ebird_code_multilabel\")\n- Only soundscape data from Zenodo formatted acoording to the Kaggle evaluation scheme.\n- Each recording is segmented into 5-second intervals where each ground truth bird vocalization is assigned to.\n- This contains segments without any labels which results in a [0] vector.\n\n**test**\n- Task: Multiclass (\"ebird_code\")\n- Only soundscape data sourced from Zenodo.\n- We provide the full recording with the complete label set and specified bounding boxes.\n- This dataset excludes recordings that do not contain bird calls (\"no_call\").","metadata":{}},{"cell_type":"markdown","source":"## Birdset Pipeline\n\n- We prepared a complete [pipeline repo](https://github.com/DBD-research-group/BirdSet/tree/main) (and a work-in-progress package) that you can work with that could ease the processing efforts.\n- You can pip install the package or clone/fork the repo. A quick tutoral notebook is [here](https://github.com/DBD-research-group/BirdSet/blob/main/notebooks/tutorials/birdset-pipeline_tutorial.ipynb).\n- Right now, we would suggest to directly work with the repo :)","metadata":{}},{"cell_type":"markdown","source":"### Download and Explore\n- You only have to download the respective dataset once, then it is handled with the HF caching\n- You get a dataset dictionary object from HF that can be treated like a pytorch dataset","metadata":{}},{"cell_type":"code","source":"import datasets\nimport transformers\n\nfrom datasets import Audio\nimport soundfile as sf\n\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2024-04-05T22:27:41.797448Z","iopub.execute_input":"2024-04-05T22:27:41.798152Z","iopub.status.idle":"2024-04-05T22:27:41.806035Z","shell.execute_reply.started":"2024-04-05T22:27:41.798102Z","shell.execute_reply":"2024-04-05T22:27:41.804276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# example dataset HSN\nHSN = datasets.load_dataset(name=\"HSN\", path='DBD-research-group/BirdSet', cache_dir='/kaggle/working')","metadata":{"execution":{"iopub.status.busy":"2024-04-05T13:14:46.489756Z","iopub.execute_input":"2024-04-05T13:14:46.490830Z","iopub.status.idle":"2024-04-05T13:14:47.382722Z","shell.execute_reply.started":"2024-04-05T13:14:46.490789Z","shell.execute_reply":"2024-04-05T13:14:47.381276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HSN","metadata":{"execution":{"iopub.status.busy":"2024-04-05T13:14:12.565449Z","iopub.execute_input":"2024-04-05T13:14:12.565847Z","iopub.status.idle":"2024-04-05T13:14:12.574562Z","shell.execute_reply.started":"2024-04-05T13:14:12.565818Z","shell.execute_reply":"2024-04-05T13:14:12.573283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The recordings are not decoded automatically. The metadata is constant across datasets but not all columns are required for training and testing respectively. \nHSN[\"train\"][0]","metadata":{"execution":{"iopub.status.busy":"2024-04-05T13:30:07.107903Z","iopub.execute_input":"2024-04-05T13:30:07.108339Z","iopub.status.idle":"2024-04-05T13:30:07.120014Z","shell.execute_reply.started":"2024-04-05T13:30:07.108303Z","shell.execute_reply":"2024-04-05T13:30:07.118578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the label in training is [ebird_code]\nprint(HSN[\"train\"].features[\"ebird_code\"])","metadata":{"execution":{"iopub.status.busy":"2024-04-05T13:29:54.309452Z","iopub.execute_input":"2024-04-05T13:29:54.311368Z","iopub.status.idle":"2024-04-05T13:29:54.323430Z","shell.execute_reply.started":"2024-04-05T13:29:54.311125Z","shell.execute_reply":"2024-04-05T13:29:54.322126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we have the ground truth labels in [ebird_code_multilabel]; no detected events required\nHSN[\"test_5s\"][0]","metadata":{"execution":{"iopub.status.busy":"2024-04-05T13:24:18.802335Z","iopub.execute_input":"2024-04-05T13:24:18.802711Z","iopub.status.idle":"2024-04-05T13:24:18.812789Z","shell.execute_reply.started":"2024-04-05T13:24:18.802682Z","shell.execute_reply":"2024-04-05T13:24:18.811476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Decoding with Huggingface\n- with cast_column, you can turn on decoding every time you access the dataset dict\n- this can be problematic since you would always decode the complete recording even if you only want to use a small slice\n- you could use e.g. a custom soundfile decoder where you only load a given segment of the file. \n- the test data test_5s is already divided into 5-second segment per sample. ","metadata":{}},{"cell_type":"code","source":"HSN = HSN.cast_column(\"audio\", Audio(sampling_rate=32_000))","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:52:06.223718Z","iopub.execute_input":"2024-04-05T14:52:06.224629Z","iopub.status.idle":"2024-04-05T14:52:06.262537Z","shell.execute_reply.started":"2024-04-05T14:52:06.224581Z","shell.execute_reply":"2024-04-05T14:52:06.261488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HSN[\"train\"][0][\"audio\"][\"array\"]","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:18:54.208365Z","iopub.execute_input":"2024-04-05T14:18:54.208812Z","iopub.status.idle":"2024-04-05T14:18:54.254355Z","shell.execute_reply.started":"2024-04-05T14:18:54.208779Z","shell.execute_reply":"2024-04-05T14:18:54.253237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(HSN[\"train\"][0][\"audio\"][\"array\"])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:19:13.995965Z","iopub.execute_input":"2024-04-05T14:19:13.996647Z","iopub.status.idle":"2024-04-05T14:19:14.391969Z","shell.execute_reply.started":"2024-04-05T14:19:13.996607Z","shell.execute_reply":"2024-04-05T14:19:14.390694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(HSN[\"test_5s\"][0][\"audio\"][\"array\"])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:19:44.908806Z","iopub.execute_input":"2024-04-05T14:19:44.909853Z","iopub.status.idle":"2024-04-05T14:19:45.274441Z","shell.execute_reply.started":"2024-04-05T14:19:44.909800Z","shell.execute_reply":"2024-04-05T14:19:45.273147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Decoding with soundfile","metadata":{}},{"cell_type":"code","source":"# example for recording with events \n\nwaveform = HSN[\"train\"][42][\"audio\"][\"array\"]\ndetected_events = HSN[\"train\"][42][\"detected_events\"]\nevent_clusters = HSN[\"train\"][42][\"event_cluster\"]\n\nunique_clusters = np.unique(event_clusters)\ncolors = plt.cm.jet(np.linspace(0, 1, len(unique_clusters))) \ncolor_map = dict(zip(unique_clusters, colors))\n\nplt.figure(figsize=(10, 4))\nplt.plot(waveform, label='Waveform')\nplt.xlabel('Sample Index')\nplt.ylabel('Amplitude')\n\ny = min(waveform)  \nheight = max(waveform) - min(waveform) \nfor (start, end), cluster in zip(detected_events, event_clusters):\n    width = end - start  \n    rect = patches.Rectangle((start*32000, y), width*32000, height, linewidth=1, \n                             edgecolor='none', facecolor=color_map[cluster], alpha=0.5)\n    plt.gca().add_patch(rect)\n\nplt.legend()\nplt.title('Audio Waveform with Detected Events')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:52:09.020578Z","iopub.execute_input":"2024-04-05T14:52:09.020998Z","iopub.status.idle":"2024-04-05T14:52:09.678095Z","shell.execute_reply.started":"2024-04-05T14:52:09.020964Z","shell.execute_reply":"2024-04-05T14:52:09.676879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# turn off decoding again\n\nHSN= HSN.cast_column(\n            column=\"audio\",\n            feature=Audio(\n                sampling_rate=32_000,\n                mono=True,\n                decode=False,\n            ),\n        )","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:41:53.613296Z","iopub.execute_input":"2024-04-05T14:41:53.613864Z","iopub.status.idle":"2024-04-05T14:41:53.651804Z","shell.execute_reply.started":"2024-04-05T14:41:53.613822Z","shell.execute_reply":"2024-04-05T14:41:53.650664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HSN[\"train\"][42][\"detected_events\"]","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:44:16.025756Z","iopub.execute_input":"2024-04-05T14:44:16.026459Z","iopub.status.idle":"2024-04-05T14:44:16.033808Z","shell.execute_reply.started":"2024-04-05T14:44:16.026423Z","shell.execute_reply":"2024-04-05T14:44:16.032677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# detected event\n\nstart = int(HSN[\"train\"][42][\"detected_events\"][0][0] * 32_000)\nend = int(HSN[\"train\"][42][\"detected_events\"][0][1] * 32_000)\naudio, sr = sf.read(HSN[\"train\"][42][\"audio\"][\"path\"], start=start, stop=end)\nplt.plot(audio)\nplt.show","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:49:39.238253Z","iopub.execute_input":"2024-04-05T14:49:39.239229Z","iopub.status.idle":"2024-04-05T14:49:39.651899Z","shell.execute_reply.started":"2024-04-05T14:49:39.239175Z","shell.execute_reply":"2024-04-05T14:49:39.650986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# complete recording \n\naudio, sr = sf.read(HSN[\"train\"][42][\"audio\"][\"path\"], start=0, stop=int(HSN[\"train\"][42][\"length\"] * 32_000))\nplt.plot(audio)\nplt.show","metadata":{"execution":{"iopub.status.busy":"2024-04-05T14:50:15.637553Z","iopub.execute_input":"2024-04-05T14:50:15.637981Z","iopub.status.idle":"2024-04-05T14:50:16.061978Z","shell.execute_reply.started":"2024-04-05T14:50:15.637945Z","shell.execute_reply":"2024-04-05T14:50:16.060890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Download only testdata","metadata":{}},{"cell_type":"code","source":"HSN = datasets.load_dataset(name=\"HSN_scape\", path='DBD-research-group/BirdSet', cache_dir='/kaggle/working')","metadata":{"execution":{"iopub.status.busy":"2024-04-05T22:27:51.629262Z","iopub.execute_input":"2024-04-05T22:27:51.629783Z","iopub.status.idle":"2024-04-05T22:29:43.045811Z","shell.execute_reply.started":"2024-04-05T22:27:51.629744Z","shell.execute_reply":"2024-04-05T22:29:43.044722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HSN","metadata":{"execution":{"iopub.status.busy":"2024-04-05T22:31:47.260713Z","iopub.execute_input":"2024-04-05T22:31:47.261278Z","iopub.status.idle":"2024-04-05T22:31:47.271959Z","shell.execute_reply.started":"2024-04-05T22:31:47.261236Z","shell.execute_reply":"2024-04-05T22:31:47.270339Z"},"trusted":true},"execution_count":null,"outputs":[]}]}