{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Finding Dataset Mean and Standard Deviation\n\nThis is a short and simple utility kernel that can help you standardize your model inputs. Standardizing inputs can help models learn more efficiently. However, when you want to do so for audio models, you should not be standardizing the raw inputs, but rather the Mel spectrograms that your model will be using, since standardizing audio inputs can give them a very different meaning after converted to spectrograms.\n\nSince your model is likely loading mini-batches one at a time, finding the right parameters for standardization is non-trivial. Use this notebook to understand.\n\nIf you are using different methods or parameters for loading and creating spectrograms for your data, please replace them with the ones used here. The standardization parameters depend on your implementation. This notebook uses librosa for loading and torchaudio (kaldi) to create spectrograms.","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\n\nimport librosa\nimport torch\nimport torchaudio.compliance.kaldi as ta_kaldi","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-19T17:57:31.858168Z","iopub.execute_input":"2023-04-19T17:57:31.858668Z","iopub.status.idle":"2023-04-19T17:57:31.864956Z","shell.execute_reply.started":"2023-04-19T17:57:31.858623Z","shell.execute_reply":"2023-04-19T17:57:31.863672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Parameters for loading and creating spectrograms\nclass CFG:\n    # Loading\n    sample_rate = 16000 # This is not the native sample rate, which is 32000, but the one I use\n    duration = 10 # seconds\n    res_type = 'soxr_hq'\n    \n    # FFT settings\n    frame_length = 25\n    frame_shift = 10\n    n_mels = 128","metadata":{"execution":{"iopub.status.busy":"2023-04-19T17:57:15.856353Z","iopub.execute_input":"2023-04-19T17:57:15.856915Z","iopub.status.idle":"2023-04-19T17:57:15.867145Z","shell.execute_reply.started":"2023-04-19T17:57:15.856848Z","shell.execute_reply":"2023-04-19T17:57:15.865839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the filepaths\nBASE_PATH = '/kaggle/input/birdclef-2023'\ndf = pd.read_csv(f'{BASE_PATH}/train_metadata.csv')\ndf['filepath'] = BASE_PATH + '/train_audio/' + df.filename","metadata":{"execution":{"iopub.status.busy":"2023-04-19T17:57:15.869346Z","iopub.execute_input":"2023-04-19T17:57:15.870489Z","iopub.status.idle":"2023-04-19T17:57:15.967993Z","shell.execute_reply.started":"2023-04-19T17:57:15.870440Z","shell.execute_reply":"2023-04-19T17:57:15.966752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the next cell, replicate your preprocessing steps. Data Augmentation and anything that does not happen in test time is not neccesary.","metadata":{}},{"cell_type":"code","source":"\nmean = []\nstd = []\nfor filepath in tqdm(df['filepath']):\n    waveform, sr = waveform, sr = librosa.load(filepath, sr=CFG.sample_rate, duration = CFG.duration, \n                                    mono=True, res_type=CFG.res_type)\n    waveform = torch.Tensor(waveform)\n    waveform = waveform.unsqueeze(0) * 2 ** 15\n    spectrogram = ta_kaldi.fbank(waveform, num_mel_bins=CFG.n_mels, sample_frequency=CFG.sample_rate, frame_length=CFG.frame_length, frame_shift=CFG.frame_shift)\n    spectrogram = spectrogram.numpy()\n    mean.append(np.mean(spectrogram))\n    std.append(np.std(spectrogram))","metadata":{"execution":{"iopub.status.busy":"2023-04-19T17:57:33.971830Z","iopub.execute_input":"2023-04-19T17:57:33.972307Z","iopub.status.idle":"2023-04-19T18:07:34.976195Z","shell.execute_reply.started":"2023-04-19T17:57:33.972265Z","shell.execute_reply":"2023-04-19T18:07:34.974749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean = np.mean(mean)\nstd = np.mean(std)\nprint(\"Mean:\", mean)\nprint(\"Standard Deviation:\", std)","metadata":{"execution":{"iopub.status.busy":"2023-04-19T18:07:55.254737Z","iopub.execute_input":"2023-04-19T18:07:55.256233Z","iopub.status.idle":"2023-04-19T18:07:55.267352Z","shell.execute_reply.started":"2023-04-19T18:07:55.256166Z","shell.execute_reply":"2023-04-19T18:07:55.265879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Take the numbers from the previous cell and put them as a parameter in your model. You only need to do this process once, unless you significantly change your preprocessing pipeline.","metadata":{}}]}