{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on March 11, 2025. By Prata, Marília (mpwolke)","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objs as go\nimport plotly.offline as py\nimport plotly.express as px\n\n#Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:03:33.230580Z","iopub.execute_input":"2025-03-12T00:03:33.230936Z","iopub.status.idle":"2025-03-12T00:03:35.499019Z","shell.execute_reply.started":"2025-03-12T00:03:33.230904Z","shell.execute_reply":"2025-03-12T00:03:35.498079Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Human Voice among the recordings\n\nOnce I heard a human voice on the 1st sample that I tried, I was intrigued about how to identify which ogg files would come from animals/birds and which were the authors registering the files that they collected.\n\nI tryied to read Librosa Notebooks, Transform Fourier, Denoising, however I didn't find any that could satisfies my search.\n\nI found on Librosa-Galery a mention about:\n\n**Music/Voice Separation Using the Similarity Matrix**\n\nAuthors:Zafar RAFII and Bryan PARDO\n\n \"In this work, we have proposed a generalization of the REpeating Pattern **Extraction Technique (REPET) method** for the task of music/voice separation, based on the calculation of a similarity matrix. The **REPET**approach is based on the separation of a musical background from a vocal foreground, by extraction of the underlying repeating structure.\"\n \n \"The basic idea is to identify elements that exhibit similar, and compare them to repeating models derived from them to extract the repeating patterns.\"\n\n \"The proposed generalization of **REPET** is only based on a **similarity matrix**. In other words, it does not depend on particular features, does not rely on complex frameworks, and does not need prior training. Because it is only based onself-similarity, it has the advantage of being simple, fast,\nblind, and therefore completely and easily automatable.\"\n\nhttps://users.cs.northwestern.edu/~zra446/doc/Rafii-Pardo%20-%20Music-Voice%20Separation%20using%20the%20Similarity%20Matrix%20-%20ISMIR%202012.pdf","metadata":{}},{"cell_type":"code","source":"# Code source: Brian McFee\n# License: ISC\n\n##################\n# Standard imports\nfrom __future__ import print_function\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport librosa\n\nimport librosa.display","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:03:45.767788Z","iopub.execute_input":"2025-03-12T00:03:45.768333Z","iopub.status.idle":"2025-03-12T00:03:45.801488Z","shell.execute_reply.started":"2025-03-12T00:03:45.768299Z","shell.execute_reply":"2025-03-12T00:03:45.800677Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Code source: Brian McFee\n# License: ISC\n\ny, sr = librosa.load('../input/birdclef-2025/train_audio/1139490/CSA36389.ogg', duration=120)\n\n\n# And compute the spectrogram magnitude and phase\nS_full, phase = librosa.magphase(librosa.stft(y))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:03:50.417705Z","iopub.execute_input":"2025-03-12T00:03:50.418039Z","iopub.status.idle":"2025-03-12T00:04:06.208972Z","shell.execute_reply.started":"2025-03-12T00:03:50.418015Z","shell.execute_reply":"2025-03-12T00:04:06.208062Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Code source: Brian McFee\n# License: ISC\n\nidx = slice(*librosa.time_to_frames([30, 35], sr=sr))\nplt.figure(figsize=(12, 4))\nlibrosa.display.specshow(librosa.amplitude_to_db(S_full[:, idx], ref=np.max),\n                         y_axis='log', x_axis='time', sr=sr)\nplt.colorbar()\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:04:11.570617Z","iopub.execute_input":"2025-03-12T00:04:11.571203Z","iopub.status.idle":"2025-03-12T00:04:12.342580Z","shell.execute_reply.started":"2025-03-12T00:04:11.571171Z","shell.execute_reply":"2025-03-12T00:04:12.341559Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Code source: Brian McFee\n# License: ISC\n\n# We'll compare frames using cosine similarity, and aggregate similar frames\n# by taking their (per-frequency) median value.\n#\n# To avoid being biased by local continuity, we constrain similar frames to be\n# separated by at least 2 seconds.\n#\n# This suppresses sparse/non-repetetitive deviations from the average spectrum,\n# and works well to discard vocal elements.\n\nS_filter = librosa.decompose.nn_filter(S_full,\n                                       aggregate=np.median,\n                                       metric='cosine',\n                                       width=int(librosa.time_to_frames(2, sr=sr)))\n\n# The output of the filter shouldn't be greater than the input\n# if we assume signals are additive.  Taking the pointwise minimium\n# with the input spectrum forces this.\nS_filter = np.minimum(S_full, S_filter)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:04:22.165067Z","iopub.execute_input":"2025-03-12T00:04:22.165386Z","iopub.status.idle":"2025-03-12T00:04:40.664611Z","shell.execute_reply.started":"2025-03-12T00:04:22.165361Z","shell.execute_reply":"2025-03-12T00:04:40.663604Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Code source: Brian McFee\n# License: ISC\n\n# We can also use a margin to reduce bleed between the vocals and instrumentation masks.\n# Note: the margins need not be equal for foreground and background separation\nmargin_i, margin_v = 2, 10\npower = 2\n\nmask_i = librosa.util.softmask(S_filter,\n                               margin_i * (S_full - S_filter),\n                               power=power)\n\nmask_v = librosa.util.softmask(S_full - S_filter,\n                               margin_v * S_filter,\n                               power=power)\n\n# Once we have the masks, simply multiply them with the input spectrum\n# to separate the components\n\nS_foreground = mask_v * S_full\nS_background = mask_i * S_full","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:04:53.747258Z","iopub.execute_input":"2025-03-12T00:04:53.747643Z","iopub.status.idle":"2025-03-12T00:04:54.067676Z","shell.execute_reply.started":"2025-03-12T00:04:53.747611Z","shell.execute_reply":"2025-03-12T00:04:54.066774Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Code source: Brian McFee\n# License: ISC\n\n# sphinx_gallery_thumbnail_number = 2\n\nplt.figure(figsize=(12, 8))\nplt.subplot(3, 1, 1)\nlibrosa.display.specshow(librosa.amplitude_to_db(S_full[:, idx], ref=np.max),\n                         y_axis='log', sr=sr)\nplt.title('Full spectrum')\nplt.colorbar()\n\nplt.subplot(3, 1, 2)\nlibrosa.display.specshow(librosa.amplitude_to_db(S_background[:, idx], ref=np.max),\n                         y_axis='log', sr=sr)\nplt.title('Background')\nplt.colorbar()\nplt.subplot(3, 1, 3)\nlibrosa.display.specshow(librosa.amplitude_to_db(S_foreground[:, idx], ref=np.max),\n                         y_axis='log', x_axis='time', sr=sr)\nplt.title('Foreground')\nplt.colorbar()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:05:01.177663Z","iopub.execute_input":"2025-03-12T00:05:01.177987Z","iopub.status.idle":"2025-03-12T00:05:02.567841Z","shell.execute_reply.started":"2025-03-12T00:05:01.177962Z","shell.execute_reply":"2025-03-12T00:05:02.566642Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Code below is wav. No output for ogg.\n\nUnfortunately, No Output below. It was a wav snippet.","metadata":{}},{"cell_type":"code","source":"#https://stackoverflow.com/questions/36458214/split-speech-audio-file-on-words-in-python\n\nfrom pydub import AudioSegment\nfrom pydub.silence import split_on_silence\n\nsound_file = AudioSegment.from_ogg(\"../input/birdclef-2025/train_audio/1139490/CSA36389.ogg\")\naudio_chunks = split_on_silence(sound_file, \n    # must be silent for at least half a second\n    min_silence_len=500,\n\n    # consider it silent if quieter than -16 dBFS\n    silence_thresh=-16\n)\n\nfor i, chunk in enumerate(audio_chunks):\n\n    out_file = \".//splitAudio//chunk{0}.ogg\".format(i)\n    print (\"exporting\"), out_file\n    chunk.export(out_file, format=\"ogg\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:05:12.205756Z","iopub.execute_input":"2025-03-12T00:05:12.206086Z","iopub.status.idle":"2025-03-12T00:05:15.992267Z","shell.execute_reply.started":"2025-03-12T00:05:12.206060Z","shell.execute_reply":"2025-03-12T00:05:15.991332Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Audio/spectogram ","metadata":{}},{"cell_type":"code","source":"#Code by Sayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook\n\nimport tensorflow as tf","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:05:26.489989Z","iopub.execute_input":"2025-03-12T00:05:26.490312Z","iopub.status.idle":"2025-03-12T00:05:40.847300Z","shell.execute_reply.started":"2025-03-12T00:05:26.490287Z","shell.execute_reply":"2025-03-12T00:05:40.846234Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Code by Sayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook\n\n!pip install noisereduce","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:05:50.680639Z","iopub.execute_input":"2025-03-12T00:05:50.681387Z","iopub.status.idle":"2025-03-12T00:05:56.205971Z","shell.execute_reply.started":"2025-03-12T00:05:50.681355Z","shell.execute_reply":"2025-03-12T00:05:56.204794Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### One audio sample.","metadata":{}},{"cell_type":"code","source":"#Code by Sayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook\n\nimport IPython\nIPython.display.Audio(\"../input/birdclef-2025/train_audio/cocwoo1/XC11212.ogg\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:07.546759Z","iopub.execute_input":"2025-03-12T00:06:07.547114Z","iopub.status.idle":"2025-03-12T00:06:07.569024Z","shell.execute_reply.started":"2025-03-12T00:06:07.547084Z","shell.execute_reply":"2025-03-12T00:06:07.567984Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Code by Sayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook\n\n!pip install tensorflow-io","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:13.205752Z","iopub.execute_input":"2025-03-12T00:06:13.206081Z","iopub.status.idle":"2025-03-12T00:06:17.284814Z","shell.execute_reply.started":"2025-03-12T00:06:13.206056Z","shell.execute_reply":"2025-03-12T00:06:17.283676Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Code by Sayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook\n\nimport tensorflow_io as tfio\nimport tensorflow as tf","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:27.251094Z","iopub.execute_input":"2025-03-12T00:06:27.251459Z","iopub.status.idle":"2025-03-12T00:06:27.715783Z","shell.execute_reply.started":"2025-03-12T00:06:27.251418Z","shell.execute_reply":"2025-03-12T00:06:27.714859Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Convert Audio to Frequency","metadata":{}},{"cell_type":"code","source":"import soundfile as sf\nfreq,rate=sf.read('../input/birdclef-2025/train_audio/ywcpar/XC115515.ogg')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:34.559004Z","iopub.execute_input":"2025-03-12T00:06:34.559332Z","iopub.status.idle":"2025-03-12T00:06:34.626659Z","shell.execute_reply.started":"2025-03-12T00:06:34.559307Z","shell.execute_reply":"2025-03-12T00:06:34.625826Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import noisereduce as nr\nreduced_noise=nr.reduce_noise(y=freq,sr=rate)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:41.418351Z","iopub.execute_input":"2025-03-12T00:06:41.418747Z","iopub.status.idle":"2025-03-12T00:06:45.949455Z","shell.execute_reply.started":"2025-03-12T00:06:41.418719Z","shell.execute_reply":"2025-03-12T00:06:45.948602Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Another sample.","metadata":{}},{"cell_type":"code","source":"import IPython\nIPython.display.Audio(\"../input/birdclef-2025/train_audio/ywcpar/XC115515.ogg\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:50.460196Z","iopub.execute_input":"2025-03-12T00:06:50.460584Z","iopub.status.idle":"2025-03-12T00:06:50.477942Z","shell.execute_reply.started":"2025-03-12T00:06:50.460551Z","shell.execute_reply":"2025-03-12T00:06:50.476799Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Convert an ogg file to spectogram","metadata":{}},{"cell_type":"code","source":"#Code by Sayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook\n\ndef read(pth,reduce_noise):\n    freq,rate=sf.read(pth)\n    if reduce_noise:\n        freq=nr.reduce_noise(y=freq,sr=rate)\n    # Convert to spectrogram\n    spectrogram = tfio.audio.spectrogram(\n    reduced_noise, nfft=3600, window=256, stride=256)\n    return tf.math.log(spectrogram).numpy()\nplt.imshow(read('../input/birdclef-2025/train_audio/ywcpar/XC115515.ogg',True));","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-12T00:06:57.658712Z","iopub.execute_input":"2025-03-12T00:06:57.659084Z","iopub.status.idle":"2025-03-12T00:07:00.393809Z","shell.execute_reply.started":"2025-03-12T00:06:57.659052Z","shell.execute_reply":"2025-03-12T00:07:00.392788Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Anyway, what should I say about this spectogram above?\n\nThe vertical axis displays frequency in Hertz, the horizontal axis represents time (just like the waveform display), and amplitude is represented by brightness. The darker background is silence?","metadata":{}},{"cell_type":"markdown","source":"### Nothing helped to clean human comments about recording species.\n\nMaybe human beings are among the Mamals ¯\\_(ツ)_/¯. Who knows: D\n\nI will wait to read the Models identifying the species of this BirdCLEF 2025.\n\nDraft Session 2h:35m","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\n Brian McFee https://librosa.org/librosa_gallery/auto_examples/plot_vocal_separation.html\n\nAnil_M https://stackoverflow.com/questions/36458214/split-speech-audio-file-on-words-in-python\n\nSayantan Mazumdar https://www.kaggle.com/swaralipibose/converting-audio-to-spectogram-noise-image-data/notebook","metadata":{}}]}