{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30683,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hello all,\n\nAs spec. + 2D CNN are poplular in BirdCLEF, faster conversion of audio files to spectrograms is helpful to accelerate experiments. The common method to convert audio files to melspectrograms is based on `librosa` or `SciPy`. In this notebook, I want to introduce a faster method based on `CuPy`, which has been validated to be effective in another competition, [HMS-HBAC](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview), and helped me a lot.\n\n**ATTENTION I**: As only CPU Notebook is allowed in inference, `CuPy` based conversion cannot be used in submission.\n\n**ATTENTION II**: `CuPy` will fail with `P100`. Therefore, make sure using `T4` in this notebook. For more detailed discussion, please refer to this [issue](https://github.com/cupy/cupy/issues/8260) in CuPy's repo.\n\nThis notebook is based on @MARK WIJKHUIZEN's [notebook](https://www.kaggle.com/code/markwijkhuizen/birdclef-2024-eda-preprocessed-dataset) and the [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487110) of @SERGEY SAHAROVSKIY in [HMS-HBAC](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview)\n\nHope this notebook is helpful for you.\n\n## Update\n\n- Update pre-processing to avoid NaN.\n- I preprocessed the training data via CuPy. You can find it in this [dataset](https://www.kaggle.com/datasets/zijiangyang1116/birdclef24-spectrograms-via-cupy).","metadata":{}},{"cell_type":"markdown","source":"# Packages","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport librosa\nfrom tqdm.notebook import tqdm","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:00.693876Z","iopub.execute_input":"2024-04-10T13:55:00.694629Z","iopub.status.idle":"2024-04-10T13:55:01.120084Z","shell.execute_reply.started":"2024-04-10T13:55:00.694593Z","shell.execute_reply":"2024-04-10T13:55:01.118985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scipy import signal as sci_signal\nimport cupy as cp\nfrom cupyx.scipy import signal as cupy_signal","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:01.122299Z","iopub.execute_input":"2024-04-10T13:55:01.123235Z","iopub.status.idle":"2024-04-10T13:55:02.450704Z","shell.execute_reply.started":"2024-04-10T13:55:01.123181Z","shell.execute_reply":"2024-04-10T13:55:02.449635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration","metadata":{}},{"cell_type":"code","source":"class Config():\n    \n    # Sample Rate\n    FS = 32000\n    \n    # make sure the spec. for each 5s audio data is [512, 512]\n    N_FFT = 1095  # N FFT\n    WIN_SIZE = 412  # WIN Size\n    WIN_LAP = 100  # OVERLAP\n    # min frequency\n    MIN_FREQ = 40\n    # max frequency\n    MAX_FREQ = 15000\n    # Competition Root Folder\n    ROOT_FOLDER = '/kaggle/input/birdclef-2024'\n    \nCONFIG = Config()","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:02.452180Z","iopub.execute_input":"2024-04-10T13:55:02.452651Z","iopub.status.idle":"2024-04-10T13:55:02.458275Z","shell.execute_reply.started":"2024-04-10T13:55:02.452620Z","shell.execute_reply":"2024-04-10T13:55:02.457208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Metadata","metadata":{}},{"cell_type":"code","source":"train_metadata_df = pd.read_csv(\n        '/kaggle/input/birdclef-2024/train_metadata.csv',\n        dtype={\n            'secondary_labels': 'string',\n            'primary_label': 'category',\n        },\n    )\n\n# Convert secondary_labels to iterable tuple\ndef parse_secondary_labels(s):\n    s = s.strip(\"[']\")\n    s = s.split(\"', '\")\n    return tuple([e for e in s if len(e) > 0])\n\ntrain_metadata_df['secondary_labels'] = train_metadata_df['secondary_labels'].apply(parse_secondary_labels)\n\n# Number of samples\nCONFIG.N_SAMPLES = len(train_metadata_df)\nprint(f'# Samples: {CONFIG.N_SAMPLES:,}')\n\ndisplay(train_metadata_df.head())\ndisplay(train_metadata_df.info())","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:02.461185Z","iopub.execute_input":"2024-04-10T13:55:02.461531Z","iopub.status.idle":"2024-04-10T13:55:02.664168Z","shell.execute_reply.started":"2024-04-10T13:55:02.461498Z","shell.execute_reply":"2024-04-10T13:55:02.663215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conversion","metadata":{}},{"cell_type":"code","source":"def oog2spec_via_cupy(audio_data):\n    audio_data = cp.array(audio_data)\n    \n    # handles NaNs\n    mean_signal = cp.nanmean(audio_data)\n    audio_data = cp.nan_to_num(audio_data, nan=mean_signal) if cp.isnan(audio_data).mean() < 1 else cp.zeros_like(audio_data)\n    \n    # to spec.\n    frequencies, times, spec_data = cupy_signal.spectrogram(\n        audio_data, \n        fs=CONFIG.FS, \n        nfft=CONFIG.N_FFT, \n        nperseg=CONFIG.WIN_SIZE, \n        noverlap=CONFIG.WIN_LAP, \n        window='hann'\n    )\n    \n    # Filter frequency range\n    valid_freq = (frequencies >= CONFIG.MIN_FREQ) & (frequencies <= CONFIG.MAX_FREQ)\n    spec_data = spec_data[valid_freq, :]\n    \n    # Log\n    spec_data = cp.log10(spec_data + 1e-20)\n    \n    # min/max normalize\n    spec_data = spec_data - spec_data.min()\n    spec_data = spec_data / spec_data.max()\n    \n    return spec_data.get()","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:02.665573Z","iopub.execute_input":"2024-04-10T13:55:02.665985Z","iopub.status.idle":"2024-04-10T13:55:02.674326Z","shell.execute_reply.started":"2024-04-10T13:55:02.665946Z","shell.execute_reply":"2024-04-10T13:55:02.673109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def oog2spec_via_scipy(audio_data):\n    # handles NaNs\n    mean_signal = np.nanmean(audio_data)\n    audio_data = np.nan_to_num(audio_data, nan=mean_signal) if np.isnan(audio_data).mean() < 1 else np.zeros_like(audio_data)\n    \n    # to spec.\n    frequencies, times, spec_data = sci_signal.spectrogram(\n        audio_data, \n        fs=CONFIG.FS, \n        nfft=CONFIG.N_FFT, \n        nperseg=CONFIG.WIN_SIZE, \n        noverlap=CONFIG.WIN_LAP, \n        window='hann'\n    )\n    \n    # Filter frequency range\n    valid_freq = (frequencies >= CONFIG.MIN_FREQ) & (frequencies <= CONFIG.MAX_FREQ)\n    spec_data = spec_data[valid_freq, :]\n    \n    # Log\n    spec_data = np.log10(spec_data + 1e-20)\n    \n    # min/max normalize\n    spec_data = spec_data - spec_data.min()\n    spec_data = spec_data / spec_data.max()\n    \n    return spec_data","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:02.675496Z","iopub.execute_input":"2024-04-10T13:55:02.675783Z","iopub.status.idle":"2024-04-10T13:55:02.684208Z","shell.execute_reply.started":"2024-04-10T13:55:02.675758Z","shell.execute_reply":"2024-04-10T13:55:02.683322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check of Consistency\n\nmake sure CuPy and Scipy generate same spectrograms.","metadata":{}},{"cell_type":"code","source":"test_audio_row = train_metadata_df.iloc[0]\ntest_file = f'{CONFIG.ROOT_FOLDER}/train_audio/{test_audio_row.filename}'\nprint(test_file)\n\n# load file\naudio_data, sample_rate = librosa.load(test_file, sr=CONFIG.FS)\nprint(audio_data.shape, sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:02.685324Z","iopub.execute_input":"2024-04-10T13:55:02.685641Z","iopub.status.idle":"2024-04-10T13:55:04.329870Z","shell.execute_reply.started":"2024-04-10T13:55:02.685610Z","shell.execute_reply":"2024-04-10T13:55:04.328897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# covert to spec. via CuPy\nspec_cupy = oog2spec_via_cupy(audio_data)\n# cover to spec. via scipy\nspec_scipy = oog2spec_via_scipy(audio_data)\n\nprint(np.mean(spec_cupy - spec_scipy))","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:04.331131Z","iopub.execute_input":"2024-04-10T13:55:04.331657Z","iopub.status.idle":"2024-04-10T13:55:05.118232Z","shell.execute_reply.started":"2024-04-10T13:55:04.331626Z","shell.execute_reply":"2024-04-10T13:55:05.117258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Very small differences","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,12))\nplt.subplot(2, 1, 1)\nplt.imshow(spec_cupy, cmap='jet')\nplt.title('Generate Spec. with CuPy')\nplt.subplot(2, 1, 2)\nplt.imshow(spec_scipy, cmap='jet')\nplt.title('Generate Spec. with SciPy')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:05.119604Z","iopub.execute_input":"2024-04-10T13:55:05.119895Z","iopub.status.idle":"2024-04-10T13:55:06.033511Z","shell.execute_reply.started":"2024-04-10T13:55:05.119870Z","shell.execute_reply":"2024-04-10T13:55:06.032584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"CuPy and Scipy can generate consistent spectrograms. Therefore, CuPy can be used locally to accelerate experiments, while Scipy can be used for inference to ensure consistency of preprocessing.","metadata":{}},{"cell_type":"markdown","source":"## Comparision of Speed","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:41:14.224597Z","iopub.execute_input":"2024-04-10T13:41:14.225102Z","iopub.status.idle":"2024-04-10T13:41:14.229589Z","shell.execute_reply.started":"2024-04-10T13:41:14.225069Z","shell.execute_reply":"2024-04-10T13:41:14.228657Z"}}},{"cell_type":"code","source":"import time\nfrom tqdm import tqdm\n\nn_max = 100\n\nbegin_time = time.time()\nfor i in tqdm(range(n_max)):\n    spec_cupy = oog2spec_via_cupy(audio_data)\nend_time = time.time()\n\nprint(f'CuPy: {(end_time - begin_time) / n_max:.4f} sec/sample')\n\nbegin_time = time.time()\nfor i in tqdm(range(n_max)):\n    spec_scipy = oog2spec_via_scipy(audio_data)\nend_time = time.time()\n\nprint(f'SciPy: {(end_time - begin_time) / n_max:.4f} sec/sample')","metadata":{"execution":{"iopub.status.busy":"2024-04-10T13:55:06.036302Z","iopub.execute_input":"2024-04-10T13:55:06.036674Z","iopub.status.idle":"2024-04-10T13:55:13.036668Z","shell.execute_reply.started":"2024-04-10T13:55:06.036635Z","shell.execute_reply":"2024-04-10T13:55:13.035706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A very significant improvement!!! CuPy is x11 faster than Scipy with Nvidia T4!!!","metadata":{}}]}