{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Audio Data Processing and EDA\n\nIn this notebook, we will start by converting the audio data into spectrograms. Spectrograms provide a visual representation of the frequency spectrum of the audio signal over time. This transformation allows us to capture the patterns and features within the audio recordings that are crucial for classifying the species based on their calls. By converting the audio files into spectrograms and saving them in the NPZ format, we ensure that the data is stored efficiently for subsequent training.\n\nAfter processing the audio data, we will perform **Exploratory Data Analysis (EDA)**. This step is important to:\n\n- Understand the distribution of species across the dataset.\n- Analyze the quality of the recordings.\n- Investigate the relationships between various metadata attributes (e.g., location, rating, etc.).\n\nBy performing EDA, we can identify any data imbalances, noisy recordings, or other factors that might affect model performance. This will help us prepare the data more effectively for the training phase.\n\n------------","metadata":{}},{"cell_type":"markdown","source":"## 🔗 BirdCLEF 2025 - Project Notebook Links\n\nHere are the different stages of my BirdCLEF 2025 pipeline, organized by functionality:\n\n### 📊 Data Preparation\n- [BirdCLEF 2025 - Data Preparation](https://www.kaggle.com/code/sheemamasood/birdclef-2025-data-prepartion)\n\n### 🎛️ Mel Spectrogram Generation\n- [BirdCLEF 2025 - Mel Generation](https://www.kaggle.com/code/sheemamasood/birdclef2025-mel-generation)\n\n### 🏷️ Pseudo Labelling for SSL\n- [BirdCLEF 2025 - Pseudo Labelling for SSL](https://www.kaggle.com/code/sheemamasood/birdclef2025-psedolabelling-for-ssl)\n\n### 🧠 Model Training\n- [BirdCLEF 2025 - Model Training (Phase 1)](https://www.kaggle.com/code/sheemamasood/birdclef2025-model-training-phase1)\n\n### 📦 Inference & Submissions\n- [BirdCLEF 2025 - Submissions](https://www.kaggle.com/code/sheemamasood/birdclef2025-submissions)\n","metadata":{}},{"cell_type":"code","source":"# Basic utilities\nimport os\nimport math\nimport time\nimport random\nimport gc\nimport logging\nimport warnings\nfrom pathlib import Path\nimport librosa.display\nimport matplotlib.pyplot as plt\n# Data handling\n\nimport numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import roc_auc_score\n\n# Audio processing\nimport librosa\n\n# Visualization\nimport cv2\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm.notebook import tqdm  # for progress bars\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Configurations:","metadata":{}},{"cell_type":"code","source":"class Config:\n    # Define paths for the dataset\n    OUTPUT_DIR = '/kaggle/working/'\n    DATA_ROOT = '/kaggle/input/birdclef-2025'  # Path to the dataset\n\n    # Audio settings\n    FS = 32000  # Sampling rate (audio)\n\n    # Mel spectrogram parameters (for converting audio to image)\n    N_FFT = 1024       # FFT window size\n    HOP_LENGTH = 512   # Step size for each frame\n    N_MELS = 128       # Number of mel bands\n    FMIN = 50          # Minimum Mel frequency\n    FMAX = 14000       # Maximum Mel frequency\n\n    # Parameters for audio duration and spectrogram size\n    TARGET_DURATION = 5.0  # Length of each audio (in seconds)\n    TARGET_SHAPE = (256, 256)  # Size of the spectrogram image\n\n    # No limit on the number of samples during training (full dataset)\n    N_MAX = None  \n\n    # flag for training mode\n    TRAINING_MODE = True  \n    \n    # Additional training-specific configurations\n    EPOCHS = 10  \n    BATCH_SIZE = 32  \n    LEARNING_RATE = 0.001  \n\n# Create the config object\nconfig = Config()\n\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Loading Data","metadata":{}},{"cell_type":"code","source":"# Load taxonomy data (bird species details)\ntaxonomy_df = pd.read_csv(f'{config.DATA_ROOT}/taxonomy.csv')\nprint(\"taxonomy data loaded\")\n\n# Create mapping from bird ID to class name\nspecies_class_map = dict(zip(taxonomy_df['primary_label'], taxonomy_df['class_name']))\n\n# Load training metadata\ntrain_df = pd.read_csv(f'{config.DATA_ROOT}/train.csv')\nprint(\"training metadata loaded \")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Basic Statistical EDA","metadata":{}},{"cell_type":"code","source":"print(\"=\"*40)\nprint(f\"📦 Train Data Shape: {train_df.shape}\")\nprint(f\"📚 Taxonomy Data Shape: {taxonomy_df.shape}\")\nprint(\"=\"*40)\n\n### 🔹 1. Columns in Each File\nprint(\"\\n🔍 Columns in Train Data:\", train_df.columns.tolist())\nprint(\"🔍 Columns in Taxonomy Data:\", taxonomy_df.columns.tolist())\n\n### 🔹 2. Data Types\nprint(\"\\n📊 Train Data Types:\")\nprint(train_df.info())\n\nprint(\"\\n📊 Taxonomy Data Types:\")\nprint(taxonomy_df.info())\n\n### 🔹 3. Basic Descriptive Statistics\nprint(\"\\n📈 Basic Stats - Train Data\")\ndisplay(train_df.describe(include='all').T)  # works well in Jupyter\n\nprint(\"\\n📈 Basic Stats - Taxonomy Data\")\ndisplay(taxonomy_df.describe(include='all').T)\n\n### 🔹 4. Missing Values Check\nprint(\"\\n❌ Missing Values in Train Data:\")\nprint(train_df.isnull().sum())\n\nprint(\"\\n❌ Missing Values in Taxonomy Data:\")\nprint(taxonomy_df.isnull().sum())\n\n### 🔹 5. Random Sample Rows for Quick Glance\nprint(\"\\n🔹 Sample Rows from Train Data:\")\ndisplay(train_df.sample(5))\n\nprint(\"\\n🔹 Sample Rows from Taxonomy Data:\")\ndisplay(taxonomy_df.sample(5))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Merging the data","metadata":{}},{"cell_type":"code","source":"# Unique bird species ke labels ko list me conversion\nlabel_list = sorted(train_df['primary_label'].unique())  # sorting unique labels\nlabel_id_list = list(range(len(label_list)))  # Har label ka ek ID number banaya\n\n# Dictionary banayi: label se id aur id se label\nlabel2id = dict(zip(label_list, label_id_list))\nid2label = dict(zip(label_id_list, label_list))\n\nprint(f'Found {len(label_list)} unique species')  # Total species print ki\n\n# Training data ka kaam karne ke liye naya dataframe banaya\nworking_df = train_df[['primary_label', 'rating', 'filename']].copy()\n\n# Har label ko uski ID \nworking_df['target'] = working_df.primary_label.map(label2id)\n\n# File ka full path\nworking_df['filepath'] = config.DATA_ROOT + '/train_audio/' + working_df.filename\n\n# Sample name banaya: foldername-filename (extension ke bina)\nworking_df['samplename'] = working_df.filename.map(lambda x: x.split('/')[0] + '-' + x.split('/')[-1].split('.')[0])\n\n# Har primary label se uski class name \nworking_df['class'] = working_df.primary_label.map(lambda x: species_class_map.get(x, 'Unknown'))\n\n# Sirf itne samples process karne hain jitne DEBUG mode ke liye allowed hain\ntotal_samples = min(len(working_df), config.N_MAX or len(working_df))\n\nprint(f'Total samples to process: {total_samples} out of {len(working_df)} available')\n\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"=\"*50)\nprint(f\"✅ Total samples to process: {total_samples} out of {len(working_df)} available\")\nprint(\"=\"*50)\n\n# 🔢 Class-wise sample count\nprint(\"\\n📊 Samples per Bird Class:\")\nprint(working_df['class'].value_counts().to_string())\n\n# 🧾 Basic Info of DataFrame\nprint(\"\\nℹ️ DataFrame Info:\")\nworking_df.info()\n\n# 📈 Basic Statistics (numerical columns like rating, target)\nprint(\"\\n📊 Descriptive Statistics:\")\nprint(working_df.describe().T)\n\n# 🔍 Missing Values Check\nprint(\"\\n❌ Missing Values Summary:\")\nprint(working_df.isnull().sum())\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Class Distribution","metadata":{}},{"cell_type":"code","source":"# Class-wise samples count ko plot karte hain\nplt.figure(figsize=(10, 6))\nsns.countplot(data=working_df, x='class', order=working_df['class'].value_counts().index)\n\n# Plot ki customization (optional)\nplt.title('Sample Distribution by Class')\nplt.xlabel('Class')\nplt.ylabel('Number of Samples')\nplt.xticks(rotation=45, ha='right')\nplt.tight_layout()\nplt.savefig('Sample Distribution by Class.png')\n# Plot \nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## No of samples per label (per specie)","metadata":{}},{"cell_type":"code","source":"# Count total unique classes\nprint(\"Total classes:\", working_df['primary_label'].nunique())\n\n# Count samples per class\nprint(\"Samples per class:\")\nprint(working_df['primary_label'].value_counts().describe())\n\n# Visualize class imbalance\nplt.figure(figsize=(10, 6))\nworking_df['primary_label'].value_counts().plot(kind='hist', bins=50, edgecolor='black')\nplt.title('Distribution of Sample Counts per Bird Species')\nplt.xlabel('Number of Samples per specie')\nplt.ylabel('Number of Classes')\nplt.tight_layout()\n\n# Save the plot\nplt.savefig('class_distribution_histogram.png')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## What This Means:\n**📉 Highly Imbalanced Dataset:**\n- Majority of the classes have very few samples (likely <100).\n- A small number of classes have hundreds of samples.\n\n**⚠️ Long Tail Problem:**\n- This is a classic long-tail distribution, common in wildlife sound datasets.\n- It means: few species are over-represented; many are under-represented.\n\n","metadata":{}},{"cell_type":"markdown","source":"###  Rating EDA working_df","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.countplot(x='rating', data=working_df)\nplt.title('Distribution of Ratings')\nplt.xlabel('Rating')\nplt.ylabel('Count')\nplt.grid(False)\nplt.savefig('Distribution of Ratings.png')\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Average and Median Rating","metadata":{}},{"cell_type":"code","source":"print(f\"Average rating: {working_df['rating'].mean():.2f}\")\nprint(f\"Median rating: {working_df['rating'].median()}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Rating by Class ","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(14, 6))\nsns.boxplot(x='class', y='rating', data=working_df)\nplt.title('Rating Distribution by Bird Class')\nplt.xlabel('Bird Class')\nplt.ylabel('Rating')\nplt.xticks(rotation=90)\nplt.tight_layout()\nplt.savefig('Rating Distribution by Bird Class.png')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📊 Basic EDA Insights\n\n- **Total Samples**: 28,564, with the majority of samples belonging to the `Aves` class (27,648), and much fewer in the other classes (`Amphibia`, `Mammalia`, and `Insecta`).\n\n- **Ratings Distribution**: Ratings range from `0.0` to `5.0`, with a mean rating of `2.92` and a standard deviation of `1.96`, indicating varied sample quality. There are no missing values in any of the columns, ensuring a clean dataset.\n","metadata":{}},{"cell_type":"markdown","source":"## Distribution of Target Labels","metadata":{}},{"cell_type":"code","source":"# Count primary labels in working_df\nprimary_label_counts = working_df['primary_label'].value_counts()\n\n# Plot distribution of all primary labels (horizontal bar plot)\nplt.figure(figsize=(15, 30))  # Adjusting the figure size to fit all labels\nax = sns.barplot(y=primary_label_counts.index, x=primary_label_counts.values)\nplt.title('Distribution of All Primary Labels (Horizontal)', fontsize=16)\nplt.ylabel('Primary Label', fontsize=14)\nplt.xlabel('Count', fontsize=14)\nplt.tight_layout()\nplt.show()\n\n# Show statistics about class imbalance\nprint(f\"Most common species: {primary_label_counts.index[0]} with {primary_label_counts.values[0]} samples\")\nprint(f\"Least common species: {primary_label_counts.index[-1]} with {primary_label_counts.values[-1]} samples\")\nprint(f\"Imbalance ratio (most common / least common): {primary_label_counts.values[0] / primary_label_counts.values[-1]:.2f}\")\n\n# Analyze the long tail\nplt.figure(figsize=(12, 6))\nplt.plot(range(len(primary_label_counts)), sorted(primary_label_counts.values, reverse=True))\nplt.title('Species Sample Count (Sorted)', fontsize=16)\nplt.xlabel('Species Rank', fontsize=14)\nplt.ylabel('Number of Samples', fontsize=14)\nplt.grid(True)\nplt.show()\n\n# Determine rare classes (< 10 samples)\nrare_classes = primary_label_counts[primary_label_counts < 10]\nprint(f\"Number of rare classes (<10 samples): {len(rare_classes)}\")\nprint(f\"Rare classes: {rare_classes.to_dict()}\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"display(working_df)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Secondary labels ","metadata":{}},{"cell_type":"code","source":"import ast\n\n# Step 1: Define function to parse string to list\ndef parse_secondary_labels(label_str):\n    if pd.isna(label_str):\n        return []\n    try:\n        return ast.literal_eval(label_str)\n    except:\n        return []\n\n# Step 2: Parse 'secondary_labels' column in train_df\ntrain_df['parsed_secondary_labels'] = train_df['secondary_labels'].apply(parse_secondary_labels)\n\n# Step 3: Merge parsed secondary labels into working_df\nworking_df = working_df.merge(\n    train_df[['filename', 'parsed_secondary_labels']],\n    on='filename',\n    how='left'\n)\n\n# Step 4: Rename for clarity\nworking_df.rename(columns={'parsed_secondary_labels': 'secondary_labels'}, inplace=True)\n\n# Step 5: Create unique ID mapping for secondary labels\nsecondary_label_list = sorted(set([label for sublist in working_df['secondary_labels'] for label in sublist]))\nsecondary_label2id = {label: idx for idx, label in enumerate(secondary_label_list)}\nid2secondary_label = {idx: label for label, idx in secondary_label2id.items()}\n\n# Step 6: Map secondary labels to ID targets\nworking_df['secondary_target'] = working_df['secondary_labels'].apply(\n    lambda x: [secondary_label2id.get(label, -1) for label in x]\n)\n\n# Step 7: Check stats\nhas_secondary = working_df['secondary_labels'].apply(lambda x: len(x) > 0)\nprint(f\"\\n✅ Recordings with secondary labels: {has_secondary.sum()} ({has_secondary.sum()/len(working_df)*100:.2f}%)\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.countplot(x=working_df['secondary_labels'].apply(len))\nplt.title('Number of Secondary Labels per Recording')\nplt.xlabel('Count of Secondary Labels')\nplt.ylabel('Number of Recordings')\nplt.savefig('Number of Secondary Labels per Recording.png')\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"display(working_df.sample(1))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### AudiO Duration Analysis","metadata":{}},{"cell_type":"code","source":"# Function to get audio duration\ndef get_audio_duration(file_path, sr=32000):  # SAMPLE_RATE jo b rakhna ho\n    try:\n        audio, _ = librosa.load(file_path, sr=sr)\n        return len(audio) / sr\n    except Exception as e:\n        print(f\"Error loading {file_path}: {e}\")\n        return None\n\n# Calculate durations for all files in working_df (or sample agar zyada ho)\ndurations = []\nfilepaths = working_df['filepath'].tolist()\n\nprint(f\"Calculating durations for {len(filepaths)} audio files...\")\n\nfor fp in tqdm(filepaths):\n    duration = get_audio_duration(fp)\n    if duration is not None:\n        durations.append(duration)\n    else:\n        durations.append(np.nan)  # handle missing if error\n\n# Add durations to working_df\nworking_df['duration'] = durations\n\n# Plot duration distribution\nplt.figure(figsize=(12, 6))\nplt.hist(working_df['duration'].dropna(), bins=50, color='skyblue')\nplt.title('Distribution of Audio Durations')\nplt.xlabel('Duration (seconds)')\nplt.ylabel('Count')\nplt.savefig(\"Distribution of Audio Durations.png\")\nplt.show()\n\n# Print some stats\nprint(f\"Duration stats:\")\nprint(f\"Mean: {np.nanmean(working_df['duration']):.2f} sec\")\nprint(f\"Median: {np.nanmedian(working_df['duration']):.2f} sec\")\nprint(f\"Min: {np.nanmin(working_df['duration']):.2f} sec\")\nprint(f\"Max: {np.nanmax(working_df['duration']):.2f} sec\")\n\n# Check short and long recordings count\nshort_count = (working_df['duration'] < 1).sum()\nlong_count = (working_df['duration'] > 60).sum()\ntotal = working_df.shape[0]\n\nprint(f\"Very short recordings (<1s): {short_count} ({short_count/total*100:.2f}%)\")\nprint(f\"Long recordings (>60s): {long_count} ({long_count/total*100:.2f}%)\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### WavePlot","metadata":{}},{"cell_type":"code","source":"# Number of waveforms to plot\nnum_samples = 5\n\n# Randomly sample files from working_df\nsampled_df = working_df.sample(n=num_samples, random_state=42).reset_index(drop=True)\n\nplt.figure(figsize=(15, 3 * num_samples))\n\nfor i, row in sampled_df.iterrows():\n    file_path = row['filepath']\n    label = row['primary_label']\n    duration = row['duration']\n\n    # Load audio\n    audio, sr = librosa.load(file_path, sr=None)  # Use native sample rate\n\n    plt.subplot(num_samples, 1, i+1)\n    librosa.display.waveshow(audio, sr=sr)\n    plt.title(f\"Waveform of {label} (Duration: {duration:.2f}s)\")\n    plt.xlabel('Time (s)')\n    plt.ylabel('Amplitude')\n\nplt.tight_layout()\n\n# Save the figure before showing\nplt.savefig('waveform_samples.png', dpi=300)\n\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### convert Audio into Spectrum","metadata":{}},{"cell_type":"code","source":"# Function jo audio ko mel spectrogram me conversion\ndef audio2melspec(audio_data):\n    # if Nan usko remove krty hain\n    if np.isnan(audio_data).any():\n        mean_val = np.nanmean(audio_data)\n        audio_data = np.nan_to_num(audio_data, nan=mean_val)\n\n    # Mel spectrogram \n    mel = librosa.feature.melspectrogram(\n        y=audio_data,\n        sr=config.FS,\n        n_fft=config.N_FFT,\n        hop_length=config.HOP_LENGTH,\n        n_mels=config.N_MELS,\n        fmin=config.FMIN,\n        fmax=config.FMAX,\n        power=2.0\n    )\n\n    # Usko decibels me conversion\n    mel_db = librosa.power_to_db(mel, ref=np.max)\n\n    # Normalization\n    mel_db = (mel_db - mel_db.min()) / (mel_db.max() - mel_db.min() + 1e-8)\n\n    return mel_db\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def prepare_audio(audio, target_len):\n    #if audio choti ho to usko repeat karo\n    while len(audio) < target_len:\n        audio = np.concatenate([audio, audio])\n\n    # Center se target length ka audio\n    start = max(0, len(audio) // 2 - target_len // 2)\n    audio = audio[start:start + target_len]\n\n    # if audio bhi choti ho to padding\n    if len(audio) < target_len:\n        audio = np.pad(audio, (0, target_len - len(audio)), mode='constant')\n    \n    return audio\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"🔄 lets start audio processing...\")\nstart_time = time.time()\n\nall_bird_data = {}\nerrors = []\ntarget_len = int(config.TARGET_DURATION * config.FS)\n\nfor i, row in tqdm(working_df.iterrows(), total=total_samples):\n    if config.N_MAX and i >= config.N_MAX:\n        break\n    try:\n        # Load karo audio\n        audio, _ = librosa.load(row.filepath, sr=config.FS)\n\n        # Audio ko prepare karo\n        audio = prepare_audio(audio, target_len)\n\n        # Mel spectrogram banao\n        mel = audio2melspec(audio)\n\n        # Agar shape match nahi karta to resize karo\n        if mel.shape != config.TARGET_SHAPE:\n            mel = cv2.resize(mel, config.TARGET_SHAPE)\n\n        # Dictionary me save karo\n        all_bird_data[row.samplename] = mel.astype(np.float32)\n\n    except Exception as e:\n        print(f\"❌ Error in {row.filepath}\")\n        errors.append((row.filepath, str(e)))\n\nend_time = time.time()\n\nprint(f\"✅ Done in {end_time - start_time:.1f} seconds\")\nprint(f\"🟢 Processed: {len(all_bird_data)} files\")\nprint(f\"🔴 Failed: {len(errors)} files\")\n\nprint(\"saving the numpy file\")\n# Save the dictionary as a NumPy compressed file (.npz)\nnp.savez_compressed('all_bird_data.npz', **all_bird_data)\n\nprint(\"✅ Data saved as all_bird_data.npz\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"working_df.to_csv('working_df.csv', index=False)\nprint(\"Saved the workin df as csv\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🔍 Steps We Followed:\n\n| 🧠 Concept Mentioned             | ✅ Implementation                                       |\n|----------------------------------|-----------------------------------------------------------------|\n| 🎧 **Audio Preprocessing**       | `librosa.load`, padding if short, centering the audio clip     |\n| 🔎 **Spectrogram Generation**    | `librosa.feature.melspectrogram()` for mel-scale spectrogram   |\n| ⚙️ **Tuned FFT Parameters**      | `n_fft`, `hop_length`, `fmin`, `fmax` from `config`            |\n| 🎚️ **Amplitude Normalization**  | `librosa.power_to_db`, then normalized to [0, 1] range         |\n| 🎯 **Optimization for Accuracy** | Centering the audio & resizing spectrograms to a fixed shape  |\n","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"# Simple list to store sample data and a set to track displayed classes\nsamples = []\ndisplayed_classes = set()\n\n# Limit the number of samples to display\nmax_samples = 4\n\n# Iterate through the dataframe\nfor i, row in working_df.iterrows():\n    if len(samples) >= max_samples:  # Stop once we've selected enough samples\n        break\n\n    if row['samplename'] in all_bird_data:\n        # If class not already displayed, add sample\n        if row['class'] not in displayed_classes:\n            samples.append((row['samplename'], row['class'], row['primary_label']))\n            displayed_classes.add(row['class'])\n\n# Plotting the spectrograms\nif samples:\n    plt.figure(figsize=(16, 12))\n    \n    # Display the spectrogram for each sample\n    for i, (samplename, class_name, species) in enumerate(samples):\n        plt.subplot(2, 2, i+1)  # Create 2x2 grid\n        plt.imshow(all_bird_data[samplename], aspect='auto', origin='lower', cmap='viridis')\n        plt.title(f\"{class_name}: {species}\")\n        plt.colorbar(format='%+2.0f dB')\n    \n    plt.tight_layout()\n    plt.savefig('melspec_examples.png')  # Save plot as an image\n    plt.show()  # Display the plot\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🖼️ Insights from Spectrogram Samples by Class\n\n| **Class**        | **Observation**                                                                 | **Interpretation**                                                       | **Suggestion**                                                                 |\n|------------------|----------------------------------------------------------------------------------|--------------------------------------------------------------------------|--------------------------------------------------------------------------------|\n| 🐜 Insecta        | Energy is mostly in **lower frequencies** with scattered high-frequency spikes  | Likely due to insect wing flapping or rhythmic chirping                  | Time-frequency patterns could be useful for classification                    |\n| 🐸 Amphibia       | Strong horizontal bands (harmonics) in **mid-frequency** range                  | Indicates tonal, repetitive calls like frog croaks                       | Harmonic structure can be extracted using MFCCs or spectral contrast features |\n| 🐶 Mammalia       | Broad frequency range with **dense energy** throughout                          | Could be due to a mix of environmental noise and mammalian vocalization  | May benefit from noise filtering or background separation                     |\n| 🐦 Aves (amalkan) | Repeating vertical stripes suggesting **repetitive bird calls**                 | Likely sharp, short chirps or tweets with structured timing              | Use temporal rhythm and pitch variation features                              |\n","metadata":{}},{"cell_type":"markdown","source":"###  Spectrogram Shapes and Stats","metadata":{}},{"cell_type":"code","source":"shapes = [mel.shape for mel in all_bird_data.values()]\nunique_shapes = set(shapes)\nprint(f\"Unique shapes: {unique_shapes}\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"###  Sample Visualizations","metadata":{}},{"cell_type":"code","source":"# Create directory to save plots if it doesn't exist\nos.makedirs(\"melspec_samples\", exist_ok=True)\n\nsample_keys = random.sample(list(all_bird_data.keys()), 5)\n\nfor key in sample_keys:\n    plt.figure(figsize=(10, 4))\n    plt.imshow(all_bird_data[key], aspect='auto', origin='lower', cmap='viridis')\n    plt.title(f\"Sample: {key}\")\n    plt.colorbar()\n    plt.tight_layout()\n    \n    # Save each image with a unique filename\n    filename = f\"melspec_samples/{key.replace('/', '_').replace(':', '_')}.png\"\n    plt.savefig(filename)\n    plt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 General Insights\n\n| **Aspect**         | **Observation**                                                  | **Suggestion**                                                      |\n|--------------------|------------------------------------------------------------------|----------------------------------------------------------------------|\n| Variation          | Different samples have different frequency-energy patterns       | dataset diverse hai — good for generalization                 |\n| Clarity            | Kuch spectrograms highly structured hain (e.g., babbly1, babbly3) | Ideal for CNNs                                                      |\n| Noise              | Kuch samples noisy ya diffuse hain (e.g., compae, orwrec3)       | Try spectrogram denoising, filtering                                |\n| Frequency bands    | Kuch samples mid-to-high frequencies mein dominate kar rahe hain | Band-wise feature extraction useful ho sakta hai                    |\n| Pattern detection  | Visual patterns clearly dikh rahe hain                           | Feature-based ya image-based deep learning (CNNs) recommend         |\n","metadata":{}},{"cell_type":"markdown","source":"### Intensity Distribution Analysis","metadata":{}},{"cell_type":"code","source":"all_values = np.concatenate([mel.flatten() for mel in all_bird_data.values()])\n# Save the plot directly without specifying the filename variable\nplt.hist(all_values, bins=50, color='skyblue')\nplt.title(\"Distribution of Mel Spectrogram Values\")\nplt.xlabel(\"Value\")\nplt.ylabel(\"Frequency\")\nplt.grid(True)\n\n# Save the plot directly\nplt.savefig(\"mel_spectrogram_distribution.png\")\nplt.close()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📊 Insights from Distribution of Mel Spectrogram Values\n\n| **Observation**               | **Interpretation**                                                      | **Suggestion**                                                        |\n|-------------------------------|-------------------------------------------------------------------------|------------------------------------------------------------------------|\n| Bell-shaped curve             | Values ka distribution roughly Gaussian lagta hai (excluding 0 spike)  | Good for models like CNNs or RNNs, especially if normalized           |\n| High zero values (left spike) | A significant portion of the data is zero-valued                        | Could be due to silence, zero-padding, or background noise            |\n| Peak around 0.4–0.5           | Most non-zero values fall in mid-intensity range                        | Ideal normalization already seems applied                             |\n| Low tail beyond 0.8           | Few high-energy components                                              | Spectrograms are not dominated by loud sounds                         |\n","metadata":{}},{"cell_type":"markdown","source":"### Duration / Energy Distribution","metadata":{}},{"cell_type":"code","source":"energies = [np.mean(mel) for mel in all_bird_data.values()]\nplt.hist(energies, bins=50, color='orange')\nplt.title(\"Mean Energy of Spectrograms\")\nplt.xlabel(\"Mean Value\")\nplt.ylabel(\"Count\")\nplt.savefig(\"Mean Energy of Spectrograms.png\")\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📊 Insights from Mean Energy of Spectrograms\n\n| **Observation**                  | **Interpretation**                                                         | **Suggestion**                                                         |\n|----------------------------------|-----------------------------------------------------------------------------|-------------------------------------------------------------------------|\n| 🔔 Bell-shaped distribution       | Mean energy values are normally distributed (Gaussian-like)                | Suitable for ML models; may benefit from standardization if needed     |\n| 🟠 Left-side spike near 0         | Many spectrograms have very low or zero energy                             | Investigate silence, padding, or low-intensity segments                |\n| 🔼 Peak between 0.4 and 0.5       | Majority of spectrograms have mid-range energy                             | Indicates consistent preprocessing or real-world sound dynamics        |\n| 🔽 Sparse tail after 0.7–0.8      | Very few samples have high mean energy                                     | Suggests lack of very loud sounds; may not need log-scaling            |\n| 📉 Symmetry of distribution       | Balanced distribution without skew                                         | Implies stable data input for training purposes                        |\n","metadata":{}},{"cell_type":"markdown","source":"## ✅ Conclusion\n\nIn this notebook, we performed a detailed Exploratory Data Analysis (EDA) on the spectrogram-based audio dataset. We explored:\n\n- Class-wise spectrogram patterns and what they reveal about the audio signals.\n- Distribution of spectrogram values and their suitability for deep learning models.\n- Identified noise, silence, and structure — crucial for preprocessing and augmentation.\n\nThe dataset shows good diversity, clear frequency patterns, and consistent energy levels — all signs of a high-quality input pipeline.\n\n---\n\n## 📦 What’s Next?\n\nWe have saved the cleaned and labeled `.npz` files. In the **next notebook**, we will:\n- Extract features (e.g., MFCC, Log-Mel)\n- Apply data augmentation for generalization\n- Build and train deep learning models (e.g., CNNs or CRNNs)\n- Evaluate performance using proper metrics\n\nStay tuned for the **Training Notebook**! 🚀\n\n---\n\n#### 📘 Made with ❤️  \nIf you found this helpful, don’t forget to **like** and **share** it with others!  \nLet's keep learning and building together. 💡🔊\n\n*******************\n********************","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}