{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"6b40dab6","cell_type":"markdown","source":"# BirdCLEF 2025 Exploratory Data Analysis\n\nThis notebook analyzes the BirdCLEF dataset to understand:\n1. Average length of training audio files\n2. Minimum number of samples per species\n3. Call/song type distribution (with sunburst visualization)\n4. Audio quality ratings and their potential use for noise handling\n5. Frequency pattern analysis for species identification","metadata":{}},{"id":"522dfe0a","cell_type":"markdown","source":"## Setup and Configuration\n\nFirst, we'll import the necessary libraries and set up our configuration. We'll use:\n- `librosa` for audio processing\n- `pandas` for data manipulation\n- `matplotlib`, `seaborn`, and `plotly` for visualizations\n- `scipy` for signal processing","metadata":{}},{"id":"8c00569e","cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport librosa\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm.notebook import tqdm\nimport warnings\nfrom pathlib import Path\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom collections import Counter, defaultdict\nfrom scipy import stats\nfrom scipy.signal import periodogram\nimport gc  # For garbage collection\nimport json\nimport time\n\nwarnings.filterwarnings('ignore')\n\n# Set display options\npd.set_option('display.max_columns', None)\nsns.set(style=\"whitegrid\")\n\n# For interactive plots in notebook\n%matplotlib inline","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:56:39.052324Z","iopub.execute_input":"2025-04-05T13:56:39.052703Z","iopub.status.idle":"2025-04-05T13:56:39.061371Z","shell.execute_reply.started":"2025-04-05T13:56:39.052673Z","shell.execute_reply":"2025-04-05T13:56:39.060242Z"}},"outputs":[],"execution_count":null},{"id":"50572a3a","cell_type":"markdown","source":"## Configuration Class\n\nWe'll set up a configuration class to store all our parameters and paths. This makes it easy to adjust settings in one place.","metadata":{}},{"id":"91345b5a","cell_type":"code","source":"class Config:\n    # Debug options\n    DEBUG_MODE = False  # Set to False for full dataset processing\n    VISUALIZE = True   # Whether to create visualization plots\n    SAMPLE_SIZE = 100  # Number of files to analyze in debug mode\n    \n    # Paths\n    BASE_DIR = '/kaggle/input/birdclef-2025/'  # Base directory for dataset\n    TRAIN_AUDIO_DIR = os.path.join(BASE_DIR, 'train_audio')\n    TRAIN_SOUNDSCAPES_DIR = os.path.join(BASE_DIR, 'train_soundscapes')\n    TEST_SOUNDSCAPES_DIR = os.path.join(BASE_DIR, 'test_soundscapes')\n    OUTPUT_DIR = \"/kaggle/working/\"\n    \n    # Audio parameters\n    SR = 32000  # Sampling rate\n    \n    def __init__(self):\n        # Adjust paths for local environment\n        if not os.path.exists(self.BASE_DIR):\n            print(\"Adjusting paths for local environment...\")\n            self.BASE_DIR = \"./data\"\n            self.TRAIN_AUDIO_DIR = \"./data/train_audio\"\n            self.TRAIN_SOUNDSCAPES_DIR = \"./data/train_soundscapes\"\n            self.TEST_SOUNDSCAPES_DIR = \"./data/test_soundscapes\" \n            self.OUTPUT_DIR = \"./output\"\n            os.makedirs(self.OUTPUT_DIR, exist_ok=True)\n\n# Initialize configuration\ncfg = Config()\n\n# Print paths\nprint(f\"Base directory: {cfg.BASE_DIR}\")\nprint(f\"Train audio directory: {cfg.TRAIN_AUDIO_DIR}\")\nprint(f\"Output directory: {cfg.OUTPUT_DIR}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:56:43.478601Z","iopub.execute_input":"2025-04-05T13:56:43.479308Z","iopub.status.idle":"2025-04-05T13:56:43.489656Z","shell.execute_reply.started":"2025-04-05T13:56:43.479269Z","shell.execute_reply":"2025-04-05T13:56:43.488193Z"}},"outputs":[],"execution_count":null},{"id":"42587137","cell_type":"markdown","source":"## Utility Functions\n\nWe'll define several functions to:\n1. Load and process audio files\n2. Assess audio quality\n3. Extract frequency characteristics\n4. Parse metadata from filenames","metadata":{}},{"id":"a9109a93","cell_type":"code","source":"# Function to load audio file and return the waveform\ndef load_audio_file(file_path, sr=Config.SR):\n    \"\"\"\n    Load audio file and return the waveform\n    \"\"\"\n    try:\n        y, _ = librosa.load(file_path, sr=sr)\n        return y\n    except Exception as e:\n        print(f\"Error loading {file_path}: {e}\")\n        return None\n\n# Function to assess audio quality\ndef assess_audio_quality(y):\n    \"\"\"\n    Assess the quality of audio based on signal statistics\n    Returns a quality score between 0 and 1\n    \"\"\"\n    if y is None or len(y) == 0:\n        return 0\n    \n    # Calculate signal statistics\n    signal_std = np.std(y)\n    signal_var = np.var(y)\n    signal_rms = np.sqrt(np.mean(np.square(y)))\n    signal_pwr = np.mean(np.square(y))\n    \n    # T = std + var + rms + pwr\n    t_stat = signal_std + signal_var + signal_rms + signal_pwr\n    \n    # Normalize to 0-1 range (approximately)\n    quality_score = min(1.0, t_stat * 100)\n    \n    return quality_score\n\n# Function to extract frequency characteristics\ndef extract_frequency_characteristics(y, sr=Config.SR):\n    \"\"\"\n    Extract frequency domain characteristics from audio\n    \"\"\"\n    if y is None or len(y) == 0:\n        return None\n    \n    # Calculate power spectral density\n    freqs, psd = periodogram(y, fs=sr)\n    \n    # Find dominant frequency\n    dominant_freq = freqs[np.argmax(psd)]\n    \n    # Calculate frequency statistics (only for audible range: 20Hz-20kHz)\n    idx_range = np.where((freqs >= 20) & (freqs <= 20000))[0]\n    \n    if len(idx_range) > 0:\n        audible_freqs = freqs[idx_range]\n        audible_psd = psd[idx_range]\n        \n        # Calculate spectral centroid (weighted average of frequencies)\n        if np.sum(audible_psd) > 0:\n            spectral_centroid = np.sum(audible_freqs * audible_psd) / np.sum(audible_psd)\n        else:\n            spectral_centroid = 0\n            \n        # Calculate spectral bandwidth (spread of frequencies)\n        if np.sum(audible_psd) > 0:\n            spectral_bandwidth = np.sqrt(np.sum(((audible_freqs - spectral_centroid) ** 2) * audible_psd) / np.sum(audible_psd))\n        else:\n            spectral_bandwidth = 0\n            \n        # Calculate spectral flatness (geometric mean / arithmetic mean)\n        # This tells us how noise-like (1) vs. tonal (0) the sound is\n        audible_psd_nonzero = audible_psd[audible_psd > 0]\n        if len(audible_psd_nonzero) > 0:\n            spectral_flatness = stats.mstats.gmean(audible_psd_nonzero) / np.mean(audible_psd_nonzero)\n        else:\n            spectral_flatness = 0\n    else:\n        spectral_centroid = 0\n        spectral_bandwidth = 0\n        spectral_flatness = 0\n    \n    return {\n        'dominant_frequency': dominant_freq,\n        'spectral_centroid': spectral_centroid,\n        'spectral_bandwidth': spectral_bandwidth,\n        'spectral_flatness': spectral_flatness\n    }\n\n# Function to extract call type from filename\ndef extract_call_type_from_filename(filename):\n    \"\"\"\n    Extract call type from filename if available\n    Common naming pattern: XC######-{call_type}.ogg\n    \"\"\"\n    call_types = [\n        'song', 'call', 'alarm', 'flight', 'drum', 'display', 'duet', \n        'juvenile', 'begging', 'contact', 'excited', 'warn', 'wing',\n        'dawn', 'growl', 'chatter', 'trill', 'whistle'\n    ]\n    \n    filename_lower = filename.lower()\n    \n    for call_type in call_types:\n        if call_type in filename_lower:\n            return call_type\n    \n    # If no specific call type found\n    if 'song' in filename_lower:\n        return 'song'\n    elif 'call' in filename_lower:\n        return 'call'\n    else:\n        return 'unknown'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:56:49.086600Z","iopub.execute_input":"2025-04-05T13:56:49.086982Z","iopub.status.idle":"2025-04-05T13:56:49.102556Z","shell.execute_reply.started":"2025-04-05T13:56:49.086953Z","shell.execute_reply":"2025-04-05T13:56:49.101140Z"}},"outputs":[],"execution_count":null},{"id":"a87a35ba","cell_type":"markdown","source":"## Load Metadata\n\nFirst, let's load the metadata from the `train.csv` file. This contains important information about each audio recording, including species labels, ratings, and geographic information.","metadata":{}},{"id":"746bb85c","cell_type":"code","source":"# Try to load metadata\ndef load_metadata(cfg):\n    try:\n        metadata_path = os.path.join(cfg.BASE_DIR, 'train.csv')\n        if os.path.exists(metadata_path):\n            metadata_df = pd.read_csv(metadata_path)\n            print(f\"Loaded metadata with {len(metadata_df)} entries\")\n            return metadata_df\n        else:\n            print(\"Metadata file not found, proceeding with file system analysis\")\n            return None\n    except Exception as e:\n        print(f\"Error loading metadata: {e}\")\n        return None\n\nmetadata_df = load_metadata(cfg)\n\n# Show a sample of the metadata if available\nif metadata_df is not None:\n    print(\"\\nMetadata columns:\", metadata_df.columns.tolist())\n    print(\"\\nSample of metadata:\")\n    display(metadata_df.head())\n    \n    # Count unique species\n    unique_species = metadata_df['primary_label'].nunique()\n    print(f\"\\nNumber of unique species: {unique_species}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:56:53.126382Z","iopub.execute_input":"2025-04-05T13:56:53.126786Z","iopub.status.idle":"2025-04-05T13:56:53.369395Z","shell.execute_reply.started":"2025-04-05T13:56:53.126748Z","shell.execute_reply":"2025-04-05T13:56:53.368132Z"}},"outputs":[],"execution_count":null},{"id":"25a7371d","cell_type":"markdown","source":"## Load Taxonomy Data\n\nThe taxonomy data provides additional information about the species, including their taxonomic classification.","metadata":{}},{"id":"15d539ac","cell_type":"code","source":"# Try to load taxonomy data\ndef load_taxonomy(cfg):\n    try:\n        taxonomy_path = os.path.join(cfg.BASE_DIR, 'taxonomy.csv')\n        if os.path.exists(taxonomy_path):\n            taxonomy_df = pd.read_csv(taxonomy_path)\n            print(f\"Loaded taxonomy data with {len(taxonomy_df)} entries\")\n            return taxonomy_df\n        else:\n            print(\"Taxonomy file not found\")\n            return None\n    except Exception as e:\n        print(f\"Error loading taxonomy: {e}\")\n        return None\n\ntaxonomy_df = load_taxonomy(cfg)\n\n# Show a sample of the taxonomy if available\nif taxonomy_df is not None:\n    print(\"\\nTaxonomy columns:\", taxonomy_df.columns.tolist())\n    print(\"\\nSample of taxonomy:\")\n    display(taxonomy_df.head())\n    \n    # Count species by class\n    if 'class' in taxonomy_df.columns:\n        class_counts = taxonomy_df['class'].value_counts()\n        print(\"\\nSpecies count by class:\")\n        display(class_counts)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:57:11.012053Z","iopub.execute_input":"2025-04-05T13:57:11.012525Z","iopub.status.idle":"2025-04-05T13:57:11.035132Z","shell.execute_reply.started":"2025-04-05T13:57:11.012477Z","shell.execute_reply":"2025-04-05T13:57:11.033818Z"}},"outputs":[],"execution_count":null},{"id":"6654b58a","cell_type":"markdown","source":"## Analyze Training Audio Files\n\nNow we'll explore the training audio files to understand:\n1. The average length of audio files\n2. The distribution of samples per species\n3. Call/song type distribution","metadata":{}},{"id":"fd545912","cell_type":"code","source":"def get_species_from_directory(cfg):\n    \"\"\"Get list of species from directory structure\"\"\"\n    if os.path.exists(cfg.TRAIN_AUDIO_DIR):\n        species_list = [d for d in os.listdir(cfg.TRAIN_AUDIO_DIR) \n                      if os.path.isdir(os.path.join(cfg.TRAIN_AUDIO_DIR, d))]\n        print(f\"Found {len(species_list)} species directories\")\n        return species_list\n    else:\n        print(f\"Training audio directory not found: {cfg.TRAIN_AUDIO_DIR}\")\n        return []\n\n# Get species from metadata or directory structure\nif metadata_df is not None:\n    species_list = metadata_df['primary_label'].unique().tolist()\n    print(f\"Found {len(species_list)} species in metadata\")\nelse:\n    species_list = get_species_from_directory(cfg)\n\n# Create species directory mapping\nspecies_dir_mapping = {}\nif os.path.exists(cfg.TRAIN_AUDIO_DIR):\n    for species in species_list:\n        species_dir = os.path.join(cfg.TRAIN_AUDIO_DIR, species)\n        if os.path.isdir(species_dir):\n            species_dir_mapping[species] = species_dir\n    \n    print(f\"Mapped {len(species_dir_mapping)} species to directories\")\n\n# Calculate samples per species\nsamples_per_species = {}\nfor species, species_dir in species_dir_mapping.items():\n    audio_files = [f for f in os.listdir(species_dir) \n                 if f.endswith(('.mp3', '.ogg', '.wav'))]\n    samples_per_species[species] = len(audio_files)\n\n# Convert to DataFrame for easier analysis\nsamples_df = pd.DataFrame({\n    'species': list(samples_per_species.keys()),\n    'sample_count': list(samples_per_species.values())\n}).sort_values('sample_count', ascending=False)\n\n# Print statistics\nprint(\"\\nSamples per species statistics:\")\nprint(f\"- Min samples: {samples_df['sample_count'].min()}\")\nprint(f\"- Max samples: {samples_df['sample_count'].max()}\")\nprint(f\"- Average samples: {samples_df['sample_count'].mean():.1f}\")\nprint(f\"- Median samples: {samples_df['sample_count'].median():.1f}\")\nprint(f\"- Total samples: {samples_df['sample_count'].sum()}\")\n\n# Show species with fewest samples\nprint(\"\\nSpecies with fewest samples:\")\ndisplay(samples_df.nsmallest(10, 'sample_count'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:57:18.220261Z","iopub.execute_input":"2025-04-05T13:57:18.220674Z","iopub.status.idle":"2025-04-05T13:57:19.877265Z","shell.execute_reply.started":"2025-04-05T13:57:18.220596Z","shell.execute_reply":"2025-04-05T13:57:19.876304Z"}},"outputs":[],"execution_count":null},{"id":"7ed11a5c","cell_type":"code","source":"# Plot distribution of samples per species\nplt.figure(figsize=(12, 6))\nplt.hist(samples_df['sample_count'], bins=50, edgecolor='black')\nplt.title('Distribution of Samples per Species')\nplt.xlabel('Number of Audio Samples')\nplt.ylabel('Number of Species')\nplt.grid(True, alpha=0.3)\nplt.show()\n\n# Plot top 30 species with most samples\nplt.figure(figsize=(14, 10))\ntop_species = samples_df.head(30)\nsns.barplot(x='sample_count', y='species', data=top_species)\nplt.title('Top 30 Species by Number of Samples')\nplt.xlabel('Number of Audio Files')\nplt.grid(True, alpha=0.3)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:57:22.415275Z","iopub.execute_input":"2025-04-05T13:57:22.415620Z","iopub.status.idle":"2025-04-05T13:57:23.654025Z","shell.execute_reply.started":"2025-04-05T13:57:22.415594Z","shell.execute_reply":"2025-04-05T13:57:23.652903Z"}},"outputs":[],"execution_count":null},{"id":"f84d5268","cell_type":"markdown","source":"## Audio Length Analysis\n\nNow we'll analyze a sample of audio files to determine their average length and distribution.","metadata":{}},{"id":"11bd6fdf","cell_type":"code","source":"def analyze_audio_length(species_list, species_dir_mapping, cfg):\n    \"\"\"Analyze the length of audio files\"\"\"\n    results = {\n        'audio_lengths': [],\n        'species_lengths': {}\n    }\n    \n    # Select a subset of species for analysis if in debug mode\n    if cfg.DEBUG_MODE:\n        np.random.seed(42)\n        if len(species_list) > 20:  # Analyze up to 20 species in debug mode\n            species_subset = np.random.choice(species_list, 20, replace=False).tolist()\n        else:\n            species_subset = species_list\n    else:\n        species_subset = species_list\n    \n    print(f\"Analyzing audio lengths for {len(species_subset)} species\")\n    \n    for species_idx, species in enumerate(species_subset):\n        # Get species directory\n        species_dir = species_dir_mapping.get(species)\n        if not species_dir or not os.path.exists(species_dir):\n            continue\n        \n        # Get audio files for this species\n        audio_files = [f for f in os.listdir(species_dir) \n                     if f.endswith(('.mp3', '.ogg', '.wav'))]\n        \n        # If in debug mode, only process a sample\n        if cfg.DEBUG_MODE:\n            if len(audio_files) > cfg.SAMPLE_SIZE:\n                audio_files = np.random.choice(audio_files, cfg.SAMPLE_SIZE, replace=False).tolist()\n        \n        species_lengths = []\n        \n        # Process audio files\n        for file_idx, filename in enumerate(tqdm(audio_files, desc=f\"Species {species_idx+1}/{len(species_subset)}\")):\n            file_path = os.path.join(species_dir, filename)\n            \n            # Load and analyze audio\n            audio = load_audio_file(file_path, sr=cfg.SR)\n            \n            if audio is not None:\n                # Calculate audio length in seconds\n                audio_length = len(audio) / cfg.SR\n                results['audio_lengths'].append(audio_length)\n                species_lengths.append(audio_length)\n        \n        # Store species-level results\n        if species_lengths:\n            results['species_lengths'][species] = species_lengths\n            \n    return results\n\n# Run the audio length analysis\naudio_length_results = analyze_audio_length(species_list, species_dir_mapping, cfg)\n\nif audio_length_results['audio_lengths']:\n    # Calculate statistics\n    audio_lengths = np.array(audio_length_results['audio_lengths'])\n    \n    print(\"\\nAudio Length Statistics:\")\n    print(f\"- Number of files analyzed: {len(audio_lengths)}\")\n    print(f\"- Average length: {np.mean(audio_lengths):.2f} seconds\")\n    print(f\"- Min length: {np.min(audio_lengths):.2f} seconds\")\n    print(f\"- Max length: {np.max(audio_lengths):.2f} seconds\")\n    print(f\"- Median length: {np.median(audio_lengths):.2f} seconds\")\n    \n    # Plot histogram of audio lengths\n    plt.figure(figsize=(12, 6))\n    plt.hist(audio_lengths, bins=50, edgecolor='black')\n    plt.axvline(np.median(audio_lengths), color='red', linestyle='dashed', \n                linewidth=1, label=f'Median: {np.median(audio_lengths):.2f}s')\n    plt.axvline(np.mean(audio_lengths), color='green', linestyle='dashed', \n                linewidth=1, label=f'Mean: {np.mean(audio_lengths):.2f}s')\n    plt.title('Distribution of Audio File Lengths')\n    plt.xlabel('Length (seconds)')\n    plt.ylabel('Count')\n    plt.grid(True, alpha=0.3)\n    plt.legend()\n    plt.show()\nelse:\n    print(\"No audio files were analyzed for length\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:57:48.921248Z","iopub.execute_input":"2025-04-05T13:57:48.921598Z","iopub.status.idle":"2025-04-05T13:59:15.520619Z","shell.execute_reply.started":"2025-04-05T13:57:48.921571Z","shell.execute_reply":"2025-04-05T13:59:15.519359Z"}},"outputs":[],"execution_count":null},{"id":"862cae87","cell_type":"markdown","source":"## Audio Length Distribution by Species\n\nLet's examine how audio lengths vary across different species.","metadata":{}},{"id":"4431c1d9","cell_type":"code","source":"# Calculate average length per species\nif audio_length_results['species_lengths']:\n    species_avg_lengths = {\n        species: np.mean(lengths)\n        for species, lengths in audio_length_results['species_lengths'].items()\n        if lengths\n    }\n    \n    # Convert to DataFrame\n    species_length_df = pd.DataFrame({\n        'species': list(species_avg_lengths.keys()),\n        'avg_length': list(species_avg_lengths.values())\n    }).sort_values('avg_length', ascending=False)\n    \n    # Display species with longest and shortest average audio lengths\n    print(\"Species with longest average audio:\")\n    display(species_length_df.head(10))\n    \n    print(\"Species with shortest average audio:\")\n    display(species_length_df.tail(10))\n    \n    # Plot average audio length by species\n    plt.figure(figsize=(14, 10))\n    sns.barplot(x='avg_length', y='species', data=species_length_df.head(20))\n    plt.title('Average Audio Length by Species (Top 20)')\n    plt.xlabel('Average Length (seconds)')\n    plt.grid(True, alpha=0.3)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:59:15.521836Z","iopub.execute_input":"2025-04-05T13:59:15.522170Z","iopub.status.idle":"2025-04-05T13:59:16.066072Z","shell.execute_reply.started":"2025-04-05T13:59:15.522141Z","shell.execute_reply":"2025-04-05T13:59:16.064952Z"}},"outputs":[],"execution_count":null},{"id":"5d778a20","cell_type":"markdown","source":"## Call Type Analysis\n\nLet's analyze the distribution of call types (song, call, alarm, etc.) in the audio files.","metadata":{}},{"id":"9076d462","cell_type":"code","source":"def analyze_call_types(species_list, species_dir_mapping, cfg):\n    \"\"\"Analyze call types from filenames\"\"\"\n    results = {\n        'call_types': {},\n        'overall_counts': Counter()\n    }\n    \n    # Select a subset of species for analysis if in debug mode\n    if cfg.DEBUG_MODE:\n        if len(species_list) > 30:  # Analyze up to 30 species in debug mode\n            species_subset = np.random.choice(species_list, 30, replace=False).tolist()\n        else:\n            species_subset = species_list\n    else:\n        species_subset = species_list\n    \n    print(f\"Analyzing call types for {len(species_subset)} species\")\n    \n    for species_idx, species in enumerate(species_subset):\n        # Get species directory\n        species_dir = species_dir_mapping.get(species)\n        if not species_dir or not os.path.exists(species_dir):\n            continue\n        \n        # Get audio files for this species\n        audio_files = [f for f in os.listdir(species_dir) \n                     if f.endswith(('.mp3', '.ogg', '.wav'))]\n        \n        species_call_types = Counter()\n        \n        # Process audio files\n        for filename in audio_files:\n            # Extract call type from filename\n            call_type = extract_call_type_from_filename(filename)\n            species_call_types[call_type] += 1\n            results['overall_counts'][call_type] += 1\n        \n        # Store species-level results\n        if species_call_types:\n            results['call_types'][species] = dict(species_call_types)\n    \n    return results\n\n# Run the call type analysis\ncall_type_results = analyze_call_types(species_list, species_dir_mapping, cfg)\n\nif call_type_results['overall_counts']:\n    # Create DataFrame for overall counts\n    call_counts = pd.DataFrame({\n        'call_type': list(call_type_results['overall_counts'].keys()),\n        'count': list(call_type_results['overall_counts'].values())\n    }).sort_values('count', ascending=False)\n    \n    print(\"\\nCall Type Distribution:\")\n    display(call_counts)\n    \n    # Plot call type distribution\n    plt.figure(figsize=(12, 8))\n    sns.barplot(x='count', y='call_type', data=call_counts)\n    plt.title('Distribution of Bird Call Types')\n    plt.xlabel('Count')\n    plt.grid(True, alpha=0.3)\n    plt.tight_layout()\n    plt.show()\n    \n    # Create pie chart\n    plt.figure(figsize=(10, 10))\n    plt.pie(call_counts['count'], labels=call_counts['call_type'], \n            autopct='%1.1f%%', startangle=90)\n    plt.axis('equal')\n    plt.title('Distribution of Bird Call Types')\n    plt.tight_layout()\n    plt.show()\nelse:\n    print(\"No call types were analyzed\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:59:16.068460Z","iopub.execute_input":"2025-04-05T13:59:16.068862Z","iopub.status.idle":"2025-04-05T13:59:16.574432Z","shell.execute_reply.started":"2025-04-05T13:59:16.068830Z","shell.execute_reply":"2025-04-05T13:59:16.572965Z"}},"outputs":[],"execution_count":null},{"id":"cc8f8333","cell_type":"markdown","source":"## Sunburst Visualization of Call Types\n\nLet's create a sunburst chart to visualize the distribution of call types across species.","metadata":{}},{"id":"ce11ae80","cell_type":"code","source":"# Prepare data for sunburst chart\nif call_type_results['call_types']:\n    # Create a list of dictionaries for the sunburst\n    sunburst_data = []\n    \n    for species, call_types in call_type_results['call_types'].items():\n        for call_type, count in call_types.items():\n            sunburst_data.append({\n                'species': species,\n                'call_type': call_type,\n                'count': count\n            })\n    \n    # Convert to DataFrame\n    sunburst_df = pd.DataFrame(sunburst_data)\n    \n    # Create sunburst chart\n    fig = px.sunburst(\n        sunburst_df,\n        path=['call_type', 'species'],\n        values='count',\n        title='Sunburst Chart of Call Types by Species',\n        width=800,\n        height=800\n    )\n    \n    fig.update_layout(margin=dict(t=30, b=30, l=0, r=0))\n    fig.show()\nelse:\n    print(\"No call type data available for sunburst visualization\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:59:16.575739Z","iopub.execute_input":"2025-04-05T13:59:16.576067Z","iopub.status.idle":"2025-04-05T13:59:19.091176Z","shell.execute_reply.started":"2025-04-05T13:59:16.576032Z","shell.execute_reply":"2025-04-05T13:59:19.089847Z"}},"outputs":[],"execution_count":null},{"id":"adde8cdc","cell_type":"markdown","source":"## Audio Quality Analysis\n\nNow we'll analyze the quality of audio files to understand how we can use this information to make our model more robust to noise.\n\nWe'll analyze both the provided quality ratings from metadata and our own quality assessment based on signal statistics.","metadata":{}},{"id":"a7235cba","cell_type":"code","source":"def analyze_quality_ratings(metadata_df):\n    \"\"\"Analyze quality ratings from metadata\"\"\"\n    if metadata_df is None or 'rating' not in metadata_df.columns:\n        print(\"No quality ratings available in metadata\")\n        return None\n    \n    # Calculate statistics\n    quality_stats = metadata_df['rating'].describe()\n    print(\"\\nQuality Rating Statistics:\")\n    display(quality_stats)\n    \n    # Count of each rating\n    rating_counts = metadata_df['rating'].value_counts().sort_index()\n    print(\"\\nCount of each rating value:\")\n    display(rating_counts)\n    \n    # Plot distribution of ratings\n    plt.figure(figsize=(10, 6))\n    sns.countplot(x='rating', data=metadata_df, palette='viridis')\n    plt.title('Distribution of Audio Quality Ratings')\n    plt.xlabel('Rating (0=unknown, 1=low, 5=high)')\n    plt.ylabel('Count')\n    plt.grid(True, alpha=0.3)\n    plt.show()\n    \n    # Average rating by species\n    species_ratings = metadata_df.groupby('primary_label')['rating'].agg(['mean', 'count']).reset_index()\n    species_ratings = species_ratings.sort_values('mean', ascending=False)\n    \n    print(\"\\nSpecies with highest average quality rating:\")\n    display(species_ratings.head(10))\n    \n    print(\"\\nSpecies with lowest average quality rating:\")\n    display(species_ratings.tail(10))\n    \n    return species_ratings\n\n# Run quality rating analysis if metadata is available\nspecies_ratings = None\nif metadata_df is not None:\n    species_ratings = analyze_quality_ratings(metadata_df)\n\n# Function to compute our own quality assessment on a sample of files\ndef analyze_computed_quality(species_list, species_dir_mapping, cfg):\n    \"\"\"Compute our own quality assessment on audio files\"\"\"\n    results = {\n        'quality_scores': [],\n        'species_scores': {}\n    }\n    \n    # Select a subset of species for analysis if in debug mode\n    if cfg.DEBUG_MODE:\n        if len(species_list) > 15:  # Analyze up to 15 species in debug mode\n            species_subset = np.random.choice(species_list, 15, replace=False).tolist()\n        else:\n            species_subset = species_list\n    else:\n        species_subset = species_list\n    \n    print(f\"Analyzing computed quality for {len(species_subset)} species\")\n    \n    for species_idx, species in enumerate(species_subset):\n        # Get species directory\n        species_dir = species_dir_mapping.get(species)\n        if not species_dir or not os.path.exists(species_dir):\n            continue\n        \n        # Get audio files for this species\n        audio_files = [f for f in os.listdir(species_dir) \n                     if f.endswith(('.mp3', '.ogg', '.wav'))]\n        \n        # If in debug mode, only process a sample\n        if cfg.DEBUG_MODE:\n            if len(audio_files) > cfg.SAMPLE_SIZE:\n                audio_files = np.random.choice(audio_files, cfg.SAMPLE_SIZE, replace=False).tolist()\n        \n        species_scores = []\n        \n        # Process audio files\n        for filename in tqdm(audio_files, desc=f\"Computing quality for {species}\"):\n            file_path = os.path.join(species_dir, filename)\n            \n            # Load and assess audio quality\n            audio = load_audio_file(file_path, sr=cfg.SR)\n            \n            if audio is not None:\n                # Compute quality score\n                quality = assess_audio_quality(audio)\n                results['quality_scores'].append((species, filename, quality))\n                species_scores.append(quality)\n        \n        # Store species-level results\n        if species_scores:\n            results['species_scores'][species] = species_scores\n    \n    return results\n\n# Run computed quality analysis\ncomputed_quality_results = analyze_computed_quality(species_list, species_dir_mapping, cfg)\n\nif computed_quality_results['quality_scores']:\n    # Extract quality scores\n    quality_scores = [score for _, _, score in computed_quality_results['quality_scores']]\n    \n    print(\"\\nComputed Quality Score Statistics:\")\n    print(f\"- Average quality: {np.mean(quality_scores):.4f}\")\n    print(f\"- Min quality: {np.min(quality_scores):.4f}\")\n    print(f\"- Max quality: {np.max(quality_scores):.4f}\")\n    \n    # Plot histogram of quality scores\n    plt.figure(figsize=(12, 6))\n    plt.hist(quality_scores, bins=50, edgecolor='black')\n    plt.title('Distribution of Computed Audio Quality Scores')\n    plt.xlabel('Quality Score (0-1)')\n    plt.ylabel('Count')\n    plt.grid(True, alpha=0.3)\n    plt.show()\n    \n    # Calculate average quality score per species\n    species_quality = {}\n    for species, scores in computed_quality_results['species_scores'].items():\n        if scores:\n            species_quality[species] = np.mean(scores)\n    \n    # Convert to DataFrame\n    species_quality_df = pd.DataFrame({\n        'species': list(species_quality.keys()),\n        'avg_quality': list(species_quality.values())\n    }).sort_values('avg_quality', ascending=False)\n    \n    # Plot average quality by species\n    plt.figure(figsize=(14, 10))\n    sns.barplot(x='avg_quality', y='species', data=species_quality_df)\n    plt.title('Average Computed Quality Score by Species')\n    plt.xlabel('Average Quality Score')\n    plt.grid(True, alpha=0.3)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T13:59:19.092195Z","iopub.execute_input":"2025-04-05T13:59:19.092532Z","iopub.status.idle":"2025-04-05T14:00:06.300704Z","shell.execute_reply.started":"2025-04-05T13:59:19.092504Z","shell.execute_reply":"2025-04-05T14:00:06.298831Z"}},"outputs":[],"execution_count":null},{"id":"2243535e","cell_type":"markdown","source":"## Frequency Pattern Analysis\n\nLet's analyze the frequency characteristics of different species to understand if we can use frequency patterns for better identification.","metadata":{}},{"id":"7629aa18","cell_type":"code","source":"def analyze_frequency_patterns(species_list, species_dir_mapping, cfg):\n    \"\"\"Analyze frequency patterns in audio files\"\"\"\n    results = {\n        'frequency_data': {}\n    }\n    \n    # Select a subset of species for analysis if in debug mode\n    if cfg.DEBUG_MODE:\n        if len(species_list) > 10:  # Analyze up to 10 species in debug mode\n            species_subset = np.random.choice(species_list, 10, replace=False).tolist()\n        else:\n            species_subset = species_list\n    else:\n        species_subset = species_list\n    \n    print(f\"Analyzing frequency patterns for {len(species_subset)} species\")\n    \n    for species_idx, species in enumerate(species_subset):\n        # Get species directory\n        species_dir = species_dir_mapping.get(species)\n        if not species_dir or not os.path.exists(species_dir):\n            continue\n        \n        # Get audio files for this species\n        audio_files = [f for f in os.listdir(species_dir) \n                     if f.endswith(('.mp3', '.ogg', '.wav'))]\n        \n        # If in debug mode, only process a sample\n        if cfg.DEBUG_MODE:\n            if len(audio_files) > cfg.SAMPLE_SIZE:\n                audio_files = np.random.choice(audio_files, cfg.SAMPLE_SIZE // 2, replace=False).tolist()\n        \n        species_freq_data = []\n        \n        # Process audio files\n        for filename in tqdm(audio_files, desc=f\"Analyzing frequencies for {species}\"):\n            file_path = os.path.join(species_dir, filename)\n            \n            # Load and analyze audio\n            audio = load_audio_file(file_path, sr=cfg.SR)\n            \n            if audio is not None:\n                # Extract frequency characteristics\n                freq_data = extract_frequency_characteristics(audio, sr=cfg.SR)\n                if freq_data:\n                    freq_data['species'] = species\n                    freq_data['filename'] = filename\n                    species_freq_data.append(freq_data)\n        \n        # Store species-level results\n        if species_freq_data:\n            results['frequency_data'][species] = species_freq_data\n    \n    return results\n\n# Run frequency pattern analysis\nfreq_pattern_results = analyze_frequency_patterns(species_list, species_dir_mapping, cfg)\n\nif freq_pattern_results['frequency_data']:\n    # Extract frequency characteristics for each species\n    species_freq_data = {}\n    for species, freq_list in freq_pattern_results['frequency_data'].items():\n        if freq_list:\n            species_freq_data[species] = {\n                'dominant_frequency': [item['dominant_frequency'] for item in freq_list],\n                'spectral_centroid': [item['spectral_centroid'] for item in freq_list],\n                'spectral_bandwidth': [item['spectral_bandwidth'] for item in freq_list],\n                'spectral_flatness': [item['spectral_flatness'] for item in freq_list]\n            }\n    \n    # Calculate median values for each characteristic by species\n    median_values = {}\n    for species, freq_data in species_freq_data.items():\n        median_values[species] = {\n            'dominant_frequency': np.median(freq_data['dominant_frequency']),\n            'spectral_centroid': np.median(freq_data['spectral_centroid']),\n            'spectral_bandwidth': np.median(freq_data['spectral_bandwidth']),\n            'spectral_flatness': np.median(freq_data['spectral_flatness'])\n        }\n    \n    # Convert to DataFrame\n    freq_df = pd.DataFrame.from_dict(median_values, orient='index')\n    \n    print(\"\\nFrequency Characteristics by Species:\")\n    display(freq_df.head())\n    \n    # Plot frequency characteristics\n    for characteristic in ['dominant_frequency', 'spectral_centroid', 'spectral_bandwidth', 'spectral_flatness']:\n        plt.figure(figsize=(14, 8))\n        sorted_df = freq_df.sort_values(characteristic, ascending=False)\n        plt.barh(range(len(sorted_df)), sorted_df[characteristic])\n        plt.yticks(range(len(sorted_df)), sorted_df.index)\n        plt.title(f'Median {characteristic.replace(\"_\", \" \").title()} by Species')\n        plt.xlabel('Value (Hz for frequencies)')\n        plt.grid(True, alpha=0.3)\n        plt.tight_layout()\n        plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T14:00:06.301546Z","iopub.execute_input":"2025-04-05T14:00:06.301850Z","iopub.status.idle":"2025-04-05T14:01:50.064090Z","shell.execute_reply.started":"2025-04-05T14:00:06.301824Z","shell.execute_reply":"2025-04-05T14:01:50.062923Z"}},"outputs":[],"execution_count":null},{"id":"40dc30ec","cell_type":"markdown","source":"## Time-Series Analysis for Bird Calls\n\nLet's explore if we can use time-series techniques like ARIMA to identify patterns in bird calls. This could potentially be used to make our model more robust.","metadata":{}},{"id":"29de7da0","cell_type":"code","source":"from statsmodels.tsa.arima.model import ARIMA\nimport matplotlib.pyplot as plt\n\ndef analyze_time_series(species_list, species_dir_mapping, cfg):\n    \"\"\"Analyze time series patterns in bird calls\"\"\"\n    # Select one species for detailed analysis\n    if species_list and species_dir_mapping:\n        # Try to find a species with known song patterns\n        target_species = None\n        for species in ['norspe1', 'comyel', 'rebnut', 'amerob', 'houspa']:\n            if species in species_dir_mapping:\n                target_species = species\n                break\n        \n        if target_species is None and species_list:\n            target_species = species_list[0]\n        \n        species_dir = species_dir_mapping.get(target_species)\n        if not species_dir or not os.path.exists(species_dir):\n            print(f\"Directory not found for species {target_species}\")\n            return\n        \n        print(f\"Performing time series analysis for {target_species}\")\n        \n        # Get audio files\n        audio_files = [f for f in os.listdir(species_dir) \n                     if f.endswith(('.mp3', '.ogg', '.wav'))]\n        \n        if not audio_files:\n            print(\"No audio files found\")\n            return\n        \n        # Select one file for analysis\n        filename = audio_files[0]\n        file_path = os.path.join(species_dir, filename)\n        \n        # Load audio\n        audio = load_audio_file(file_path, sr=cfg.SR)\n        \n        if audio is None:\n            print(f\"Failed to load audio file: {file_path}\")\n            return\n        \n        # Plot waveform\n        plt.figure(figsize=(14, 5))\n        plt.subplot(2, 1, 1)\n        plt.plot(np.arange(len(audio)) / cfg.SR, audio)\n        plt.title(f'Waveform for {target_species} - {filename}')\n        plt.xlabel('Time (s)')\n        plt.ylabel('Amplitude')\n        plt.grid(True, alpha=0.3)\n        \n        # Power spectral density\n        plt.subplot(2, 1, 2)\n        freqs, psd = periodogram(audio, fs=cfg.SR)\n        plt.semilogy(freqs, psd)\n        plt.title('Power Spectral Density')\n        plt.xlabel('Frequency (Hz)')\n        plt.ylabel('Power/Frequency (dB/Hz)')\n        plt.grid(True, alpha=0.3)\n        plt.tight_layout()\n        plt.show()\n        \n        # Extract envelope as a time series\n        def extract_envelope(signal, frame_size=1024, hop_length=512):\n            energy = []\n            for i in range(0, len(signal) - frame_size, hop_length):\n                chunk = signal[i:i + frame_size]\n                energy.append(np.sum(chunk ** 2) / frame_size)\n            return np.array(energy)\n        \n        # Get envelope\n        envelope = extract_envelope(audio)\n        \n        # Plot envelope\n        plt.figure(figsize=(14, 5))\n        plt.plot(np.arange(len(envelope)) * (512 / cfg.SR), envelope)\n        plt.title(f'Energy Envelope for {target_species} - {filename}')\n        plt.xlabel('Time (s)')\n        plt.ylabel('Energy')\n        plt.grid(True, alpha=0.3)\n        plt.show()\n        \n        # Try fitting an ARIMA model to the envelope\n        try:\n            # Use a small sample to make computation faster\n            sample_size = min(len(envelope), 1000)\n            sample_envelope = envelope[:sample_size]\n            \n            # Fit ARIMA model\n            model = ARIMA(sample_envelope, order=(5, 1, 0))\n            model_fit = model.fit()\n            \n            # Print summary\n            print(\"\\nARIMA Model Summary:\")\n            print(model_fit.summary())\n            \n            # Plot forecast vs actual\n            plt.figure(figsize=(14, 5))\n            plt.plot(sample_envelope, label='Actual')\n            plt.plot(model_fit.fittedvalues, color='red', label='Fitted')\n            plt.title(f'ARIMA Model Fit for {target_species}')\n            plt.legend()\n            plt.grid(True, alpha=0.3)\n            plt.show()\n            \n            # Forecast next 100 points\n            forecast = model_fit.forecast(steps=100)\n            \n            # Plot forecast\n            plt.figure(figsize=(14, 5))\n            plt.plot(range(len(sample_envelope)), sample_envelope, label='Actual')\n            plt.plot(range(len(sample_envelope), len(sample_envelope) + 100), forecast, color='red', label='Forecast')\n            plt.title(f'ARIMA Forecast for {target_species}')\n            plt.legend()\n            plt.grid(True, alpha=0.3)\n            plt.show()\n            \n        except Exception as e:\n            print(f\"Error fitting ARIMA model: {e}\")\n\n# Run time series analysis\nanalyze_time_series(species_list, species_dir_mapping, cfg)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-05T14:01:50.065059Z","iopub.execute_input":"2025-04-05T14:01:50.065348Z","iopub.status.idle":"2025-04-05T14:01:54.947579Z","shell.execute_reply.started":"2025-04-05T14:01:50.065322Z","shell.execute_reply":"2025-04-05T14:01:54.945958Z"}},"outputs":[],"execution_count":null},{"id":"7df0ad1f","cell_type":"markdown","source":"## Conclusions and Next Steps\n\nBased on our exploratory data analysis, we've discovered:\n\n1. **Audio Length Distribution**: The average length of training audio files is approximately [X] seconds, with significant variation across species. This information will help us determine optimal window sizes for processing.\n\n2. **Sample Distribution**: We identified species with very few samples, with the minimum being [X] samples. This highlights potential class imbalance issues we'll need to address.\n\n3. **Call Type Distribution**: The most common call types are [X, Y, Z], with \"song\" and \"call\" being predominant. Our sunburst visualization shows how call types vary across species.\n\n4. **Audio Quality**: We analyzed both the provided quality ratings and computed our own quality scores. This can be used to develop a noise identification model as suggested.\n\n5. **Frequency Patterns**: We identified distinct frequency characteristics for different species, which could be leveraged for better classification.\n\n### Next Steps:\n\n1. **Noise-Robust Model**: Develop a model to identify and characterize noise in audio segments, then use this to enhance training data with similar noise patterns.\n\n2. **Time-Series Features**: Implement ARIMA and other time-series techniques to extract temporal patterns in bird calls.\n\n3. **Enhanced Data Augmentation**: Use the identified call characteristics to create more targeted augmentation strategies.\n\n4. **Class Balancing**: Address the imbalanced nature of the dataset, especially for species with few samples.\n\n5. **Sliding Window Approach**: Optimize the 5-second window approach for inference based on the average call length we discovered.","metadata":{}},{"id":"2ccd3d58-04ae-45a0-ac5c-138ecae41094","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"8f49e3ca-8dd6-4751-8b9b-5c8af956bb9f","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}