{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BirdCLEF 2024 - Data Preprocessing\n\nThis notebook preprocesses the BirdCLEF 2024 dataset for training a deep \nlearning model. The dataset consists of audio files in OGG format and metadata\nin CSV format. The metadata includes the primary and secondary labels of the \naudio files. The primary label is the main bird species in the audio file, \nwhile the secondary labels are additional bird species present in the audio \nfile. The goal is to classify the primary bird species in the audio file.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport imageio.v3 as imageio\nimport albumentations as A\n\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split\n\nimport librosa\nimport cv2\nimport pickle","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:26.137972Z","iopub.execute_input":"2024-06-10T13:15:26.138388Z","iopub.status.idle":"2024-06-10T13:15:27.825244Z","shell.execute_reply.started":"2024-06-10T13:15:26.138349Z","shell.execute_reply":"2024-06-10T13:15:27.824278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Config\n\nThe `Config` class contains the configuration parameters for the preprocessing.\nThe parameters include the horizontal resolution of the mel spectrogram, the \nroot folder of the competition data, the maximum decibel to clip the audio to,\nthe minimum rating of the audio files, the sample rate, the number of samples \nin the window for the Fourier transform, and the step size of the window.","metadata":{}},{"cell_type":"code","source":"class Config():\n    # Horizontal melspectrogram resolution\n    MELSPEC_H = 128\n    # Competition Root Folder\n    ROOT_FOLDER = '/kaggle/input/birdclef-2024'\n    # Maximum decibel to clip audio to\n    TOP_DB = 100\n    # Minimum rating\n    MIN_RATING = 3.0\n    # Sample rate as provided in competition description\n    SR = 32000\n    N_FFT = 2000\n    HOP_LENGTH = 500\n\nCONFIG = Config()","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:27.827138Z","iopub.execute_input":"2024-06-10T13:15:27.827593Z","iopub.status.idle":"2024-06-10T13:15:27.833010Z","shell.execute_reply.started":"2024-06-10T13:15:27.827562Z","shell.execute_reply":"2024-06-10T13:15:27.831888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pd.read_csv(f'{CONFIG.ROOT_FOLDER}/sample_submission.csv')\n\n# Set labels\nCONFIG.LABELS = sample_submission.columns[1:]\nCONFIG.N_LABELS = len(CONFIG.LABELS)\nCONFIG.N_CLASSES = len(CONFIG.LABELS)\n","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:27.834295Z","iopub.execute_input":"2024-06-10T13:15:27.834633Z","iopub.status.idle":"2024-06-10T13:15:27.862244Z","shell.execute_reply.started":"2024-06-10T13:15:27.834606Z","shell.execute_reply":"2024-06-10T13:15:27.861069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Metadata\n\nThe metadata CSV file contains the following columns:\n- `filename`: the name of the audio file\n- `primary_label`: the primary bird species in the audio file\n- `secondary_labels`: additional bird species in the audio file\n- `rating`: the rating of the audio file\nand other columns.\n\nBut in this notebook, we only need the `filename`, `primary_label`, `ratings`, and `secondary_labels` columns.\n","metadata":{}},{"cell_type":"code","source":"train_metadata_df = pd.read_csv(\n        f'{CONFIG.ROOT_FOLDER}/train_metadata.csv',\n        dtype={\n            'secondary_labels': 'string',\n            'primary_label': 'category',\n        },\n    )\n\n# Convert secondary_labels to iterable tuple\ndef parse_secondary_labels(s):\n    s = s.strip(\"[']\")\n    s = s.split(\"', '\")\n    return tuple([e for e in s if len(e) > 0])\n\ntrain_metadata_df['secondary_labels'] = train_metadata_df['secondary_labels'].apply(parse_secondary_labels)\n\n# Number of samples\nCONFIG.N_SAMPLES = len(train_metadata_df)\nprint(f'# Samples: {CONFIG.N_SAMPLES:,}')","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:27.863603Z","iopub.execute_input":"2024-06-10T13:15:27.863949Z","iopub.status.idle":"2024-06-10T13:15:28.069119Z","shell.execute_reply.started":"2024-06-10T13:15:27.863917Z","shell.execute_reply":"2024-06-10T13:15:28.068042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function to convert OGG audio files to melspectrogram encoded as PNG bytes\n\nThe function `ogg2melspectrogram` takes the path of an OGG audio file and\nconverts it to a mel spectrogram. The function uses the `librosa` library to\nload the audio file, normalize the audio, and convert it to a mel spectrogram.\nThe mel spectrogram is then converted to decibels and normalized to 0-255. The\nspectrogram is then converted to PNG bytes using the `cv2` library.\n\nThe function returns the PNG bytes of the mel spectrogram.","metadata":{}},{"cell_type":"code","source":"# Convert OOG audio files to melspectrogram encoded as PNG bytes\ndef ogg2melspectrogram(file_path):\n    # Load the audio file\n    y, _ = librosa.load(file_path, sr=CONFIG.SR)\n    # Normalize audio\n    y = librosa.util.normalize(y)\n    # Convert to mel spectrogram\n    spec = librosa.feature.melspectrogram(\n        y=y,\n        sr=CONFIG.SR, # sample rate\n        n_fft=CONFIG.N_FFT, # number of samples in window \n        hop_length=CONFIG.HOP_LENGTH, # step size of window\n        n_mels=CONFIG.MELSPEC_H, # horizontal resolution from fmin→fmax in log scale\n        fmin=40, # minimum frequency\n        fmax=15000, # maximum frequency\n        power=2.0, # intensity^power for log scale\n    )\n    # Convert to Db\n    spec = librosa.power_to_db(spec, ref=CONFIG.TOP_DB)\n    # Normalize 0-min\n    spec = spec - spec.min()\n    # Normalize 0-255\n    spec = (spec / spec.max() * 255).astype(np.uint8)\n    # Convert to PNG bytes\n    _, spec_png_uint8 = cv2.imencode('.png', spec)\n    spec_png_bytes = bytes(spec_png_uint8)\n    \n    return spec_png_bytes","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:28.072178Z","iopub.execute_input":"2024-06-10T13:15:28.072639Z","iopub.status.idle":"2024-06-10T13:15:28.080626Z","shell.execute_reply.started":"2024-06-10T13:15:28.072601Z","shell.execute_reply":"2024-06-10T13:15:28.079325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# A bit more Image analysis\nsample_audio_file = f'{CONFIG.ROOT_FOLDER}/train_audio/commoo3/XC739282.ogg'\nsample_spectrogram = ogg2melspectrogram(sample_audio_file)\nsample_spec = imageio.imread(sample_spectrogram)\n\nfrom IPython.display import Image\ndisplay(Image(sample_spectrogram))\n\nplt.imshow(sample_spec)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:28.082083Z","iopub.execute_input":"2024-06-10T13:15:28.082472Z","iopub.status.idle":"2024-06-10T13:15:45.698763Z","shell.execute_reply.started":"2024-06-10T13:15:28.082436Z","shell.execute_reply":"2024-06-10T13:15:45.697566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing\n\nThe data preprocessing involves the following steps:\n1. Map the class labels to integer labels\n2. Filter the audio files based on the minimum rating and absence of secondary labels\n3. Convert the audio files to mel spectrograms\n","metadata":{}},{"cell_type":"code","source":"# Maps a class to corresponding integer label\nCLASS2LABEL = dict(zip(CONFIG.LABELS, np.arange(CONFIG.N_LABELS)))\n# Label to class mapping\nLABEL2CLASS = dict([(v,k) for k, v in CLASS2LABEL.items()])\n# Create dataset\nX = {}\ny = {}\nfor idx, row in tqdm(train_metadata_df.iterrows(), total=CONFIG.N_SAMPLES):\n    # Check for absence of secondary label and minimum rating\n    if row['rating'] >= CONFIG.MIN_RATING and len(row['secondary_labels']) == 0:\n        # Save PNGs in X\n        X[idx] = ogg2melspectrogram(f'{CONFIG.ROOT_FOLDER}/train_audio/{row.filename}')\n        # Save labels in y\n        y[idx] = CLASS2LABEL.get(row['primary_label'])\n        \nprint(f'# Training Samples: {len(X):,}')","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:15:45.700078Z","iopub.execute_input":"2024-06-10T13:15:45.700604Z","iopub.status.idle":"2024-06-10T13:17:56.646478Z","shell.execute_reply.started":"2024-06-10T13:15:45.700560Z","shell.execute_reply":"2024-06-10T13:17:56.644950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example Plots\ndef example_plots(X, y, N):\n    np.random.seed(45)\n    random_keys = np.random.choice(list(X.keys()), N)\n    \n    for k in random_keys:\n        spec = imageio.imread(X[k])\n        plt.figure(figsize=(12,5))\n        plt.title(\n                f'Label: {y[k]}, Class: {LABEL2CLASS[y[k]]}, shape: {spec.shape}, ' +\n                f'min: {spec.min():.0f}, max: {spec.max():.0f}, ' +\n                f'µ: {spec.mean():.1f}, σ: {spec.std():.1f}'\n            )\n        plt.imshow(spec)\n        plt.show()\n        \nexample_plots(X, y, 8)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:17:56.647520Z","iopub.status.idle":"2024-06-10T13:17:56.647905Z","shell.execute_reply.started":"2024-06-10T13:17:56.647713Z","shell.execute_reply":"2024-06-10T13:17:56.647729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split Data\n\nThe data is split into training and validation sets using the `train_test_split`\nfunction from the `sklearn` library. The training set contains 80% of the data,\nwhile the validation set contains 20% of the data. The data is stratified based\non the labels to ensure that the distribution of the labels is the same in both\nthe training and validation sets.","metadata":{}},{"cell_type":"code","source":"# Split data\n\n# Convert the dictionaries to lists\nX_list = [v for k, v in X.items()]\ny_list = [v for k, v in y.items()]\n\n# Split data into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(\n    X_list, y_list, test_size=0.2, random_state=42, stratify=y_list\n)\n\n# Convert them back to dictionaries\nX_train = {i: X_train[i] for i in range(len(X_train))}\ny_train = {i: y_train[i] for i in range(len(y_train))}\nX_val = {i: X_val[i] for i in range(len(X_val))}\ny_val = {i: y_val[i] for i in range(len(y_val))}\n\n\nprint(f'# Training Samples: {len(X_train):,}')\nprint(f'# Validation Samples: {len(X_val):,}')","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:17:56.649922Z","iopub.status.idle":"2024-06-10T13:17:56.650268Z","shell.execute_reply.started":"2024-06-10T13:17:56.650104Z","shell.execute_reply":"2024-06-10T13:17:56.650119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save Data\n\nThe training and validation data are saved as pickle files for later use.\n","metadata":{}},{"cell_type":"code","source":"# Write training data\n\nwith open(f'X.pkl', 'wb') as f:\n    pickle.dump(X, f)\n\nwith open(f'y.pkl', 'wb') as f:\n    pickle.dump(y, f)\n\nwith open(f'X_train.pkl', 'wb') as f:\n    pickle.dump(X_train, f)\n\nwith open(f'y_train.pkl', 'wb') as f:\n    pickle.dump(y_train, f)\n    \n# Write validation data\nwith open(f'X_val.pkl', 'wb') as f:\n    pickle.dump(X_val, f)\n    \nwith open(f'y_val.pkl', 'wb') as f:\n    pickle.dump(y_val, f)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T13:17:56.651546Z","iopub.status.idle":"2024-06-10T13:17:56.651888Z","shell.execute_reply.started":"2024-06-10T13:17:56.651718Z","shell.execute_reply":"2024-06-10T13:17:56.651732Z"},"trusted":true},"execution_count":null,"outputs":[]}]}