{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **BirdCLEF+ 2025 Competition - Part 1 - Data Exploration and Initial Pre-processing**","metadata":{}},{"cell_type":"markdown","source":"This notebook will primarily handle data exploration and initial pre-processing. \n\nGiven this is my first time working on a multi-label audio classification task, I'll try my level best to flesh out and structure my work as best as possible in - what I imagine will be - a series of notebooks which will cover various aspects of the data preparation, model training, iteration and submission process.\n\nAdditionally, I am testing the impact of long context LLMS on my workflows and this competition will be no exception. So my work will be augmented with the support of a mixture of LLMs as I progress through the pipeline. It should be noted that as per the competition rules, I will do my level best to abide by the \"Reasonableness\" standards.\n\nIf anyone finds this useful then please do hit the `UPVOTE` button as a show of support.\n\nNow, onto the task at hand....","metadata":{}},{"cell_type":"markdown","source":"## **Imports and Setup**","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os, gc, random \nfrom pathlib import Path\nfrom tqdm.notebook import tqdm\nimport IPython.display as ipd\nfrom IPython.display import display, clear_output\nimport ipywidgets as widgets\n\nimport librosa\nimport librosa.display\nimport soundfile as sf\n\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score, accuracy_score, confusion_matrix","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.383496Z","iopub.execute_input":"2025-04-17T09:58:24.383810Z","iopub.status.idle":"2025-04-17T09:58:24.417260Z","shell.execute_reply.started":"2025-04-17T09:58:24.383779Z","shell.execute_reply":"2025-04-17T09:58:24.416347Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class Config:\n    def __init__(self, **kwargs):\n        for k, v in kwargs.items():\n            setattr(self, k, v)\n\n    def update(self, **kwargs):\n        for k, v in kwargs.items():\n            setattr(self, k, v)\n\n# Initialize and set basic configuration\ncfg = Config(SEED=42, SAMPLE_RATE=32000,\n             DATA_PATH=Path(\"/kaggle/input/birdclef-2025\"))\n\n# Verifying changes\nprint(cfg.__dict__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.418961Z","iopub.execute_input":"2025-04-17T09:58:24.419925Z","iopub.status.idle":"2025-04-17T09:58:24.425605Z","shell.execute_reply.started":"2025-04-17T09:58:24.419892Z","shell.execute_reply":"2025-04-17T09:58:24.424941Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Function to seed everything to ensure reproducibility\ndef seed_everything(seed):\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False # Change to true if input sizes are kept constant\n\nseed_everything(cfg.SEED)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.426328Z","iopub.execute_input":"2025-04-17T09:58:24.426611Z","iopub.status.idle":"2025-04-17T09:58:24.452219Z","shell.execute_reply.started":"2025-04-17T09:58:24.426581Z","shell.execute_reply":"2025-04-17T09:58:24.451432Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Device check\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Using device: {device}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.453092Z","iopub.execute_input":"2025-04-17T09:58:24.453309Z","iopub.status.idle":"2025-04-17T09:58:24.461010Z","shell.execute_reply.started":"2025-04-17T09:58:24.453290Z","shell.execute_reply":"2025-04-17T09:58:24.460039Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Data Structure**\n\nLet's attempt to review the various file types, including the metadata. According to the competition's data tab, the file `train.csv` contains the metadata we are looking for.","metadata":{}},{"cell_type":"code","source":"# Files in the base data path\nprint(f\"Files in the base data path include: {os.listdir(cfg.DATA_PATH)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.463102Z","iopub.execute_input":"2025-04-17T09:58:24.464107Z","iopub.status.idle":"2025-04-17T09:58:24.479571Z","shell.execute_reply.started":"2025-04-17T09:58:24.464084Z","shell.execute_reply":"2025-04-17T09:58:24.478788Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Taking a closer look at the meta data\nmetadata_path = cfg.DATA_PATH / \"train.csv\"\n\nif metadata_path.exists():\n    train_df = pd.read_csv(metadata_path)\n    print(train_df.head(15))\n    print(\"\\nMetadata Columns:\", train_df.columns)\n    print(\"\\nTraining Samples:\", len(train_df))\n    print(\"\\nUnique Species:\", train_df['primary_label'].nunique())\n    print(\"\\nSecondary Species Labels(Recordist Marked):\", train_df['secondary_labels'].nunique())\n    # Key distributions\n    print(\"\\nSpecies distribution (top 10):\")\n    print(train_df[['primary_label', 'scientific_name']].value_counts().head(10))\n    print(\"\\nSpecies distribution (bottom 10):\")\n    print(train_df[['primary_label', 'scientific_name']].value_counts().tail(10))\nelse:\n    print(f\"Metadata file not found at {meta_datapath}. Check path!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T11:20:56.494872Z","iopub.execute_input":"2025-04-17T11:20:56.495972Z","iopub.status.idle":"2025-04-17T11:20:56.708956Z","shell.execute_reply.started":"2025-04-17T11:20:56.495926Z","shell.execute_reply":"2025-04-17T11:20:56.707935Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.760287Z","iopub.execute_input":"2025-04-17T09:58:24.760652Z","iopub.status.idle":"2025-04-17T09:58:24.825240Z","shell.execute_reply.started":"2025-04-17T09:58:24.760618Z","shell.execute_reply":"2025-04-17T09:58:24.824248Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Taking a closer look at the label distributions to get a better sense of class imbalances.","metadata":{}},{"cell_type":"code","source":"train_df.describe(include=[object])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.826181Z","iopub.execute_input":"2025-04-17T09:58:24.826466Z","iopub.status.idle":"2025-04-17T09:58:24.916618Z","shell.execute_reply.started":"2025-04-17T09:58:24.826445Z","shell.execute_reply":"2025-04-17T09:58:24.915781Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.describe(include=[np.number])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.917633Z","iopub.execute_input":"2025-04-17T09:58:24.917937Z","iopub.status.idle":"2025-04-17T09:58:24.946795Z","shell.execute_reply.started":"2025-04-17T09:58:24.917914Z","shell.execute_reply":"2025-04-17T09:58:24.946134Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- There are 206 classes with a significant imbalance i.e. the top species `grekis` (also known by its scientific name `Pitangus sulphuratus`) has ~2x more samples than the 10th species `wbwwre1` (Henicorhina leucosticta). This also means that the top species has a representation which is orders of magnitude more than the rarest species (see distributions in the outputs above).\n    - This strongly reinforces the need for modified loss functions designed to address class imbalances.\n- At this stage, I'll treat location data as secondary i.e. useful for context building. This can be revisited later.\n- Highly varied or _multi-taxa_ goals, as per the `taxonomy.csv` file, suggest that classification won't just be limited to birds, but will also include amphibians, mammals and insects. Thus the model architecture needs to account for this. ","metadata":{}},{"cell_type":"markdown","source":"## **Exploration and Pre-processing**","metadata":{}},{"cell_type":"markdown","source":"Key spectrogram parameters (`N_MELS`, `N_FFT`, `HOP_LENGTH`, `FMIN`, `FMAX`) will be defined in our `Config` class. These significantly impact the resulting image and model performance. Specifically:\n- `N_FFT` controls frequency resolution.\n- `HOP_LENGTH` controls time resolution.\n- `N_MELS` is the number of frequency bins in the final Mel-spectrogram.\n- `FMIN` and `FMAX` help focus on the relevant frequency range for birds.\n\nAdditionally, converting to decibels (`power_to_db`) makes faint sounds more visibile in the plots.","metadata":{}},{"cell_type":"code","source":"# Update config\ncfg.N_MELS = 128           # number of MEL bands(can be adjusted after experimentation)\ncfg.N_FFT = 2048           # window size for fast fourier transform (FFT)\ncfg.HOP_LENGTH = 512       # number of samples b/w successive frames\ncfg.FMIN = 50              # minimum frequency\ncfg.FMAX = 14000           # maximum frequency (relevant for bird calls)\n# New Clip Params\ncfg.TARGET_DURATION_S = 5  # setting at 5 secs to make it easier to hear context\ncfg.TARGET_SAMPLES = cfg.TARGET_DURATION_S * cfg.SAMPLE_RATE","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.948904Z","iopub.execute_input":"2025-04-17T09:58:24.949122Z","iopub.status.idle":"2025-04-17T09:58:24.953554Z","shell.execute_reply.started":"2025-04-17T09:58:24.949105Z","shell.execute_reply":"2025-04-17T09:58:24.952788Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Verifying changes\nprint(cfg.__dict__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.954467Z","iopub.execute_input":"2025-04-17T09:58:24.954705Z","iopub.status.idle":"2025-04-17T09:58:24.968982Z","shell.execute_reply.started":"2025-04-17T09:58:24.954687Z","shell.execute_reply":"2025-04-17T09:58:24.968305Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Helper function to plot waveform and spectrograms\ndef plot_spectrogram(waveform, sr, title=\"Waveform and Mel Spectrogram\"):\n    \"\"\"Plots waveform and Mel spectrogram for a given audio signal\"\"\"\n    fig, axs = plt.subplots(2, 1, figsize=(12, 8), sharex=True)\n\n    # Plot waveform\n    librosa.display.waveshow(waveform, sr=sr, ax=axs[0])\n    axs[0].set_title('Waveform')\n    axs[0].set_ylabel('Amplitude')\n\n    # Generate and plot Mel spectrogram\n    mel_spectrogram = librosa.feature.melspectrogram(y=waveform, sr=sr,\n                                                     n_fft=cfg.N_FFT,\n                                                     hop_length=cfg.HOP_LENGTH,\n                                                     n_mels=cfg.N_MELS,\n                                                     fmin=cfg.FMIN,\n                                                     fmax=cfg.FMAX)\n    # Using power_to_db converts amplitude spectrogram to dB scale to improve visuals\n    mel_spectrogram_db = librosa.power_to_db(mel_spectrogram, ref=np.max)\n    img = librosa.display.specshow(mel_spectrogram_db, sr=sr, hop_length=cfg.HOP_LENGTH,\n                                   x_axis='time', y_axis='mel', ax=axs[1],\n                                   fmin=cfg.FMIN, fmax=cfg.FMAX)\n    axs[1].set_title('Mel Spectrogram (dB)')\n    axs[1].set_ylabel('Mel Frequency')\n    axs[1].set_xlabel('Time (s)')\n    fig.colorbar(img, ax=axs[1], format='%+2.0f dB')\n\n    plt.suptitle(title, fontsize=17)\n    plt.tight_layout(rect=[0, 0.03, 1, 0.95]) # Adjust layout to prevent title overlap\n    plt.show()\n    return mel_spectrogram_db","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.969900Z","iopub.execute_input":"2025-04-17T09:58:24.970184Z","iopub.status.idle":"2025-04-17T09:58:24.986159Z","shell.execute_reply.started":"2025-04-17T09:58:24.970156Z","shell.execute_reply":"2025-04-17T09:58:24.985320Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load and visualize a few random samples from the training data\nif 'train_df' in globals() and not train_df.empty: # check of DF exists and isn't empty\n    num_samples = 5\n    sample_files = train_df.sample(num_samples, random_state=cfg.SEED)\n\n    for index, row in sample_files.iterrows():\n        filename = row['filename']\n        label = row['primary_label']\n        file_path = cfg.DATA_PATH/\"train_audio\"/filename\n\n        if file_path.exists():\n            print(f\"Loading: {filename}; Label: {label}\")\n            try:\n                # Load audio file and use duration=None initially to load the full file\n                waveform, sr = librosa.load(file_path, sr=cfg.SAMPLE_RATE, duration=None)\n                print(f\"    Original Duration: {len(waveform) / sr:.2f} seconds.\")\n\n                # Play 10 sec snippet\n                display_duration = min(10.0, len(waveform) / sr)\n                print(f\"   Playing first {display_duration:.2f} seconds:\")\n                ipd.display(ipd.Audio(waveform[:int(display_duration*sr)], rate=sr))\n\n                # Plot waveform and spectrogram of the first chunk (e.g. 5 seconds)\n                # Use the whole file if shorter than 5 secs\n                plot_waveform, _ = librosa.load(file_path, sr=cfg.SAMPLE_RATE, duration=cfg.TARGET_DURATION_S)\n                _ = plot_spectrogram(plot_waveform, cfg.SAMPLE_RATE, \n                                     title=f\"{filename}; (Label: {label}); First {cfg.TARGET_DURATION_S}s\")\n            \n            except Exception as e:\n                print(f\"Error processing {filename}: {e}\")\n            print(\"-\" * 30)\n        else:\n            print(f\"File not found: {file_path}\")\nelse:\n    print(\"Training metadata DataFrame ('train_df') not found or is empty! Skipping sample loading.\")\n    print(\"Try to manually specify some audio file paths for exploration.\")\n                ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:24.986918Z","iopub.execute_input":"2025-04-17T09:58:24.987174Z","iopub.status.idle":"2025-04-17T09:58:46.315007Z","shell.execute_reply.started":"2025-04-17T09:58:24.987156Z","shell.execute_reply":"2025-04-17T09:58:46.314200Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Handling Variable Audio Lengths & Segmentation**","metadata":{"execution":{"iopub.status.busy":"2025-04-15T11:52:22.289821Z","iopub.execute_input":"2025-04-15T11:52:22.290319Z","iopub.status.idle":"2025-04-15T11:52:22.294704Z","shell.execute_reply.started":"2025-04-15T11:52:22.290294Z","shell.execute_reply":"2025-04-15T11:52:22.293772Z"}}},{"cell_type":"markdown","source":"As in the case of most deep learning models, inputs need to be of a fixed size. Audio recordings are rarely of uniform length, and our training dataset is no exception.\n\nA common strategy is to break the original inputs into shorter, fixed length clips (for e.g. 3 or 5 seconds). This should take care of longer recordings while focusing the model on relevant short events. An added benefit of shortened clips is an increase in the number of training samples.","metadata":{}},{"cell_type":"code","source":"# Analyze audio durations\nif 'train_df' in globals() and not train_df.empty:\n    print(\"Analyzing audio durations...\")\n    durations = []\n    pbar = tqdm(train_df['filename'].tolist(), desc=\"Calculating durations\")\n    for filename in pbar:\n        file_path = cfg.DATA_PATH/\"train_audio\"/filename\n        if file_path.exists():\n            try:\n                # Efficient approach to get duration with loading the whole file\n                info = sf.info(file_path)\n                durations.append(info.duration)\n            except Exception as e:\n                print(f\"Could not get info for {filename}: {e}\") #Comment / uncomment for debugging\n                durations.append(np.nan) # mark errors\n        else:\n            durations.append(np.nan)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T09:58:46.315963Z","iopub.execute_input":"2025-04-17T09:58:46.316535Z","iopub.status.idle":"2025-04-17T10:03:26.559280Z","shell.execute_reply.started":"2025-04-17T09:58:46.316501Z","shell.execute_reply":"2025-04-17T10:03:26.558211Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['duration'] = durations # new column for durations\n#train_df.dropna(subset=['duration'], inplace=True) # remove rows where duration couldn't be calculated\ntrain_df['duration'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T10:03:26.560175Z","iopub.execute_input":"2025-04-17T10:03:26.560468Z","iopub.status.idle":"2025-04-17T10:03:26.575968Z","shell.execute_reply.started":"2025-04-17T10:03:26.560447Z","shell.execute_reply":"2025-04-17T10:03:26.575167Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df['duration'].isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T10:03:26.577070Z","iopub.execute_input":"2025-04-17T10:03:26.577415Z","iopub.status.idle":"2025-04-17T10:03:26.590598Z","shell.execute_reply.started":"2025-04-17T10:03:26.577365Z","shell.execute_reply":"2025-04-17T10:03:26.589659Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(12, 8))\nsns.histplot(train_df['duration'], bins=100)\nplt.title('Distribution of Audio File Durations (s)')\nplt.xlabel('Duration (s)')\nplt.ylabel('Count')\nplt.show()\n\nprint(train_df['duration'].describe())\nprint(f\"\\nNumber of clips if using {cfg.TARGET_DURATION_S}s segments (approx.):\")\n# calculate total duration / target segment duration\ntotal_segments = np.ceil(train_df['duration'] / cfg.TARGET_DURATION_S).sum()\nprint(f\"  ~ {int(total_segments):,} segments\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T10:03:26.591618Z","iopub.execute_input":"2025-04-17T10:03:26.591932Z","iopub.status.idle":"2025-04-17T10:03:26.941438Z","shell.execute_reply.started":"2025-04-17T10:03:26.591907Z","shell.execute_reply":"2025-04-17T10:03:26.940611Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Initial observations from the waveform and mel spectrograms reveals distinct patterns for each species with varying levels of background noise (i.e. noise of additional flora and fauna in wild settings).\n- Observations also have varying levels of quality ratings (Values are 1 to 5  with 1=low quality, and 5=high quality) provided by users of Xeno-canto (0 implies no rating is available; iNaturalist and the CSA do not provide quality ratings.)","metadata":{}},{"cell_type":"code","source":"# Function to create fixed length clips\ndef create_clips(waveform, sr, clip_length_samples, overlap_samples=0, end_behavior='pad'):\n    \"\"\" Splits or pads a waveform into fixed length clips\"\"\"\n    clips = []\n    total_samples = len(waveform)\n    step = clip_length_samples - overlap_samples\n    current_pos = 0\n\n    while current_pos < total_samples:\n        end_pos = current_pos + clip_length_samples\n        clip = waveform[current_pos:end_pos]\n\n        # Handle end of the waveform\n        if len(clip) < clip_length_samples:\n            if end_behavior == 'pad':\n                padding_needed = clip_length_samples - len(clip)\n                clip = np.pad(clip, (0, padding_needed), 'constant')\n                clips.append(clip)\n            elif end_behavior == 'truncate':\n                pass # discard if the remaining part is shorter than the clip length\n            elif end_behavior == 'variable':\n                # Add shorter clip - requires downstream logic to handle variable sizes or padding later\n                if len(clip) > 0: # Avoid empty clips if overlap > step\n                    clips.append(clip)\n            else:\n                raise ValueError(f\"Unknown end_behavior: {end_behavior}\")\n            break # End of waveform\n        else:\n            clips.append(clip)\n\n        if step <= 0: # Avoid infinite loop if overlap >= clip_length\n            raise ValueError(\"Overlap must be less than clip length\")\n        current_pos += step\n\n        # Ensure we don't go past the end if step makes us jumpover the last samples\n        # mostly relevant if overlap > 0\n        if current_pos >= total_samples and end_behavior != 'variable':\n            break # Prevents adding a fully paddedd clip unnecessarily when using 'pad'\n    \n    return clips\n      ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T10:03:26.942550Z","iopub.execute_input":"2025-04-17T10:03:26.942800Z","iopub.status.idle":"2025-04-17T10:03:26.950131Z","shell.execute_reply.started":"2025-04-17T10:03:26.942781Z","shell.execute_reply":"2025-04-17T10:03:26.949070Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'waveform' in globals(): # last loaded waveform \n    print(f\"\\nSplitting example waveform (duration: {len(waveform) / cfg.SAMPLE_RATE:.2f}s) into\\\n    {cfg.TARGET_DURATION_S}s clips...\")\n    example_clips = create_clips(waveform, cfg.SAMPLE_RATE, cfg.TARGET_SAMPLES, end_behavior='pad')\n    print(f\"Generated {len(example_clips)} clips.\")\n\n    # Visualizing the first clip\n    if example_clips:\n        print(\"Visualizing the first generated clip:\")\n        _ = plot_spectrogram(example_clips[0], cfg.SAMPLE_RATE, title=\"First 5s Clip\")\n    else:\n        print(\"No clips generated (original file may be too short).\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T10:03:26.951064Z","iopub.execute_input":"2025-04-17T10:03:26.951360Z","iopub.status.idle":"2025-04-17T10:03:27.814843Z","shell.execute_reply.started":"2025-04-17T10:03:26.951331Z","shell.execute_reply":"2025-04-17T10:03:27.814025Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Interactive Exploration Using `ipywidgets`** (_Viewable in a Jupyter Kernel_)","metadata":{}},{"cell_type":"markdown","source":"We can use `ipywidgets` to quickly browse through the file folders while plotting their spectrograms. \n\nI am sure that most Kagglers are familiar with `ipywidgets`, but in case anyone isn't then this [notebook](https://www.kaggle.com/code/maevaralafi/diving-into-ipython-widgets/notebook) serves as an excellent guide for the uninitiated or for those who're looking to refresh their widget building skills.","metadata":{}},{"cell_type":"code","source":"# Setting up widget functionality while modifying the code we used to generate mel-spectrograms and waveforms\nif 'train_df' in globals() and not train_df.empty:\n    # Ensure duration column exists from previous steps\n    if 'duration' not in train_df.columns:\n        print(\"Warning: `duration` column not found in train_df.\")\n        # Limiting the size of the interactive_df for performance.\n        interactive_df = train_df.sample(n=50, random_state=42) # Reduce to 50 to improve performance\n    else:\n        # Selecting files with a reasonable duration\n        interactive_df = train_df[(train_df['duration'] >= cfg.TARGET_DURATION_S)].copy()\n\n    if not interactive_df.empty:\n        # Create drop down menu items from filenames and labels\n        options = [\n            (f\"{row['filename']} ({row['primary_label']})\", idx)\n            for idx, row in interactive_df.sample(min(50, len(interactive_df)), random_state=cfg.SEED).iterrows() # Reduce to 50\n        ]\n        options.sort() # alphabetical sorting\n\n        file_dropdown = widgets.Dropdown(options=options, description='Select File:')\n        output_area = widgets.Output()\n\n        def on_file_change(change):\n            with output_area:\n                clear_output(wait=True) # Clear previous output\n                if change['new'] is not None:\n                    selected_index = change['new']\n                    row = interactive_df.loc[selected_index]\n                    filename = row['filename']\n                    label = row['primary_label']\n                    file_path = cfg.DATA_PATH/\"train_audio\"/filename\n\n                    if file_path.exists():\n                        print(f\"Loading: {filename} (Label: {label})\")\n                        try:\n                            # Increased target duration for more context\n                            waveform, sr = librosa.load(file_path, sr=cfg.SAMPLE_RATE, duration=cfg.TARGET_DURATION_S*2)\n                            display(ipd.Audio(waveform[:cfg.TARGET_SAMPLES], rate=sr)) # Play first segment\n                            _ = plot_spectrogram(waveform[:cfg.TARGET_SAMPLES], sr, \n                                                 title=f\"{filename} ({label} - First {cfg.TARGET_DURATION_S}s)\")\n                        except Exception as e:\n                            print(f\"      Error processing {filename}: {e}\")\n                    else:\n                        print(\"File not found: {file_path}\")\n\n        file_dropdown.observe(on_file_change, names='value')\n\n        print(\"Interactive Spectrogram Viewer:\")\n        display(file_dropdown)\n        display(output_area)\n        \n        # Trigger initial loadd for default selection\n        on_file_change({'new': file_dropdown.value})\n    \n    else:\n        print(\"No suitable files found for interactive exploration (duration >= 5s).\")\nelse:\n    print(\"Train metadata DataFrame ('train_df') not found or empty. Skipping interactive viewer.\")\n                            ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-17T10:03:27.815911Z","iopub.execute_input":"2025-04-17T10:03:27.816229Z","iopub.status.idle":"2025-04-17T10:03:28.752475Z","shell.execute_reply.started":"2025-04-17T10:03:27.816205Z","shell.execute_reply":"2025-04-17T10:03:28.751593Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- The mix of both clean recordings and noisy field recordings, indicating varying SNR (signal to noise ratio), appears to be quite typical. The model needs to be robust to this noise. Which brings us back to the ratings column.\n    - Ratings could potentially be used to filter out very low quality samples (e.g. rating 1) if they appear to be detrimental during training.\n    - The caveat here is that there are numerous ratings with a value of 0, so we will need to be flexible with this variable.\n- The positively skewed distribution of audio sample `durations` (ranging between 0.5s to ~20mins), necessitates the need for a clipping strategy, as shown above. A median of ~21s means a typical file will yeild ~4 clips of 5 seconds.\n    - Clipping also gives segments with varying signal characteristics. The reason for this is that we are imposing  a fixed 5 second window constraint onto a continuous and variable audio stream. So we will end up capturing either the start, middle, end or silence between various calls.\n    - Introducing this type of variability is beneficial for robustness.\n\n\nWith exploration and preprocessing covered for now, I'll proceed with setting up the preprocessing pipeline utility functions, dataloaders, and augmentation functions in the next notebook.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}