{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":19596,"databundleVersionId":1292430,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Bird Species Classification from Audio Recordings using Acoustic Feature Engineering\n\n## Objectives\n- Extract handcrafted acoustic features (MFCC, spectral descriptors, chroma, zero-crossing rate) from bird-call recordings in the Cornell Birdcall Identification (`birdsong-recognition`) dataset.\n- Visualize spectrograms and inspect what the extracted features look like for a sample recording.\n- Compare feature-selection strategies — variance threshold, ANOVA F-test, and Random Forest importance — to find the most informative acoustic features.\n- Benchmark classical ML models (Random Forest, Logistic Regression, KNN, SVM) on a manageable subset of well-represented species, reporting **accuracy and macro-F1** since classes aren't perfectly balanced.\n- Address class imbalance with **SMOTE**, and compare it against a **conditional GAN** trained to generate synthetic minority-class feature vectors.\n- Try a **GRU-based deep learning model** on raw mel-spectrogram sequences as an alternative to handcrafted features.\n\n## Requirements\n- Kaggle notebook with the **Cornell Birdcall Identification** competition dataset attached at `/kaggle/input/birdsong-recognition`.\n- `librosa` for audio processing, `scikit-learn` / `imbalanced-learn` for classical ML, and `tensorflow`/`keras` for the GRU and GAN sections (all pre-installed on Kaggle).\n","metadata":{}},{"cell_type":"markdown","source":"## 0. Setup","metadata":{}},{"cell_type":"code","source":"# Only install what's missing. Do NOT upgrade scikit-learn here: if sklearn has\n# already been imported elsewhere in this kernel session, an upgrade installs new\n# files on disk but the OLD module stays cached in memory, and mixing the two\n# causes: ImportError: cannot import name '_fit_context' from 'sklearn.base'.\n# Installing imbalanced-learn without --upgrade lets pip pick a version compatible\n# with whatever scikit-learn is already running.\n!pip install -q imbalanced-learn\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:17:42.541985Z","iopub.execute_input":"2026-07-12T09:17:42.542441Z","iopub.status.idle":"2026-07-12T09:17:46.423284Z","shell.execute_reply.started":"2026-07-12T09:17:42.542403Z","shell.execute_reply":"2026-07-12T09:17:46.421962Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport warnings\nimport numpy as np\nimport pandas as pd\nimport librosa\nimport librosa.display\nimport matplotlib.pyplot as plt\nfrom collections import Counter\n\n# The Xeno-Canto mp3s in this dataset aren't perfectly clean, so the mp3 decoder\n# (mpg123, via audioread) and librosa both like to warn about it -- \"Xing stream\n# size off by more than 1%\", \"Illegal Audio-MPEG-Header\", \"PySoundFile failed.\n# Trying audioread instead\", tuning-estimation warnings on near-silent clips, etc.\n# These are almost always harmless (librosa still recovers usable audio), so we\n# quiet the Python-level ones here. The mpg123 \"Note:\"/\"Warning:\" lines print\n# straight to the console from the C decoder and can't be silenced this way --\n# they're noise, not failures, unless followed by an actual Python exception.\nwarnings.filterwarnings(\"ignore\", category=FutureWarning, module=\"librosa\")\nwarnings.filterwarnings(\"ignore\", category=UserWarning, module=\"librosa\")\n\ninput_dir = '/kaggle/input/birdsong-recognition/'\n\n# Print a compact overview instead of every single filename (there are\n# tens of thousands of audio files -- listing them all floods the output).\nfor dirname, subdirs, filenames in os.walk(input_dir):\n    if filenames:\n        print(f\"{dirname}: {len(filenames)} files\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:17:46.425762Z","iopub.execute_input":"2026-07-12T09:17:46.426213Z","iopub.status.idle":"2026-07-12T09:18:08.961761Z","shell.execute_reply.started":"2026-07-12T09:17:46.426181Z","shell.execute_reply":"2026-07-12T09:18:08.960872Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\nprint(os.listdir(\"/kaggle/input\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:08.962882Z","iopub.execute_input":"2026-07-12T09:18:08.963535Z","iopub.status.idle":"2026-07-12T09:18:08.968470Z","shell.execute_reply.started":"2026-07-12T09:18:08.963509Z","shell.execute_reply":"2026-07-12T09:18:08.967635Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1. Explore a Sample Spectrogram","metadata":{}},{"cell_type":"code","source":"def audio_to_melspec(path, sr=16000, n_mels=64, hop=512):\n    y, _ = librosa.load(path, sr=sr)\n    mel = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=n_mels, hop_length=hop)\n    mel_db = librosa.power_to_db(mel, ref=np.max)\n    return mel_db\n\n# Pick the first available species/file dynamically instead of hardcoding a\n# filename that may not exist in every snapshot of the dataset.\ntrain_audio_dir = os.path.join(input_dir, 'train_audio')\nexample_species = sorted(os.listdir(train_audio_dir))[0]\nexample_file = sorted(os.listdir(os.path.join(train_audio_dir, example_species)))[0]\nexample_path = os.path.join(train_audio_dir, example_species, example_file)\nprint(\"Example file:\", example_path)\n\nspec = audio_to_melspec(example_path)\nplt.figure(figsize=(10, 4))\nlibrosa.display.specshow(spec, sr=16000, x_axis=\"time\", y_axis=\"mel\")\nplt.colorbar(format='%+2.0f dB')\nplt.title(f\"Mel-Spectrogram of Birdsong ({example_species})\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:08.971020Z","iopub.execute_input":"2026-07-12T09:18:08.971344Z","iopub.status.idle":"2026-07-12T09:18:09.428341Z","shell.execute_reply.started":"2026-07-12T09:18:08.971317Z","shell.execute_reply":"2026-07-12T09:18:09.427397Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Feature Extraction Functions\n\nEach recording is reduced to a **37-dimensional feature vector**: 20 MFCCs + 4 spectral descriptors + 12 chroma bins + 1 zero-crossing rate.","metadata":{}},{"cell_type":"code","source":"def extract_mfcc(y, sr, n_mfcc=20):\n    mfccs = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=n_mfcc)\n    return np.mean(mfccs, axis=1)\n\ndef extract_spectral(y, sr):\n    centroid = np.mean(librosa.feature.spectral_centroid(y=y, sr=sr))\n    bandwidth = np.mean(librosa.feature.spectral_bandwidth(y=y, sr=sr))\n    rolloff = np.mean(librosa.feature.spectral_rolloff(y=y, sr=sr))\n    contrast = np.mean(librosa.feature.spectral_contrast(y=y, sr=sr))\n    return np.array([centroid, bandwidth, rolloff, contrast])\n\ndef extract_chroma(y, sr):\n    chroma = librosa.feature.chroma_stft(y=y, sr=sr)\n    return np.mean(chroma, axis=1)\n\ndef extract_zcr(y):\n    return np.array([np.mean(librosa.feature.zero_crossing_rate(y))])\n\ndef extract_all(path, sr=16000):\n    y, sr = librosa.load(path, sr=sr)\n    return np.hstack([\n        extract_mfcc(y, sr),\n        extract_spectral(y, sr),\n        extract_chroma(y, sr),\n        extract_zcr(y)\n    ])\n\ndef safe_extract(path, sr=16000):\n    \"\"\"extract_all(), but resilient to real-world audio problems: a handful of\n    clips in this dataset are truncated/corrupted mp3s that the decoder can't\n    fully resync on (you may see mpg123 \"Giving up resync\" in the console for\n    these), and a few near-silent clips produce non-finite feature values\n    (e.g. librosa's tuning estimation on an empty frequency set). Rather than\n    crashing the whole extraction loop on one bad file, this returns None and\n    the caller skips it.\"\"\"\n    try:\n        feats = extract_all(path, sr=sr)\n    except Exception as e:\n        print(f\"Skipping (failed to decode): {path} -- {type(e).__name__}: {e}\")\n        return None\n    if not np.all(np.isfinite(feats)):\n        print(f\"Skipping (non-finite features, likely a silent/corrupt clip): {path}\")\n        return None\n    return feats\n\n# Human-readable names for the 37 columns above -- built once, reused everywhere\n# below instead of being rebuilt (and risking drift) in multiple cells.\nFEATURE_NAMES = (\n    [f\"MFCC_{i+1}\" for i in range(20)] +\n    [\"Spectral_Centroid\", \"Spectral_Bandwidth\", \"Spectral_Rolloff\", \"Spectral_Contrast\"] +\n    [f\"Chroma_{i+1}\" for i in range(12)] +\n    [\"ZeroCrossingRate\"]\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:09.429519Z","iopub.execute_input":"2026-07-12T09:18:09.429881Z","iopub.status.idle":"2026-07-12T09:18:09.442266Z","shell.execute_reply.started":"2026-07-12T09:18:09.429830Z","shell.execute_reply":"2026-07-12T09:18:09.441342Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def extract_and_plot(path):\n    \"\"\"Verbose version of extract_all(): prints each feature group and plots\n    the waveform, MFCCs, chroma, and spectral centroid/rolloff for inspection.\"\"\"\n    print(f\"\\n=== Extracting features from: {path} ===\")\n    y, sr = librosa.load(path, sr=16000)\n    print(f\"Loaded audio: {len(y)} samples at {sr} Hz\")\n\n    mfccs = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=20)\n    mfcc_mean = np.mean(mfccs, axis=1)\n    print(f\"[MFCC] shape: {mfccs.shape}, mean vector (first 5): {mfcc_mean[:5]}\")\n\n    centroid = librosa.feature.spectral_centroid(y=y, sr=sr)[0]\n    bandwidth = librosa.feature.spectral_bandwidth(y=y, sr=sr)[0]\n    rolloff = librosa.feature.spectral_rolloff(y=y, sr=sr)[0]\n    contrast = np.mean(librosa.feature.spectral_contrast(y=y, sr=sr))\n    print(f\"[Spectral] mean centroid={np.mean(centroid):.2f}, bandwidth={np.mean(bandwidth):.2f}, \"\n          f\"rolloff={np.mean(rolloff):.2f}, contrast={contrast:.2f}\")\n\n    chroma = librosa.feature.chroma_stft(y=y, sr=sr)\n    chroma_mean = np.mean(chroma, axis=1)\n    print(f\"[Chroma] shape: {chroma.shape}, mean vector (first 5): {chroma_mean[:5]}\")\n\n    zcr = np.mean(librosa.feature.zero_crossing_rate(y))\n    print(f\"[ZCR] mean zero crossing rate: {zcr:.4f}\")\n\n    all_feats = np.hstack([mfcc_mean,\n                            [np.mean(centroid), np.mean(bandwidth), np.mean(rolloff), contrast],\n                            chroma_mean, [zcr]])\n    print(f\"[All Features] Final vector shape: {all_feats.shape}\")\n\n    # --- PLOTS (this block was cut off / unclosed in the original notebook) ---\n    fig, ax = plt.subplots(2, 2, figsize=(12, 8))\n\n    librosa.display.waveshow(y, sr=sr, ax=ax[0, 0])\n    ax[0, 0].set_title(\"Waveform\")\n\n    img1 = librosa.display.specshow(mfccs, x_axis='time', ax=ax[0, 1])\n    ax[0, 1].set_title(\"MFCCs\")\n    fig.colorbar(img1, ax=ax[0, 1])\n\n    img2 = librosa.display.specshow(chroma, y_axis='chroma', x_axis='time', ax=ax[1, 0])\n    ax[1, 0].set_title(\"Chroma Features\")\n    fig.colorbar(img2, ax=ax[1, 0])\n\n    ax[1, 1].plot(centroid, label=\"Spectral Centroid\")\n    ax[1, 1].plot(rolloff, label=\"Spectral Rolloff\", alpha=0.7)\n    ax[1, 1].set_title(\"Spectral Centroid & Rolloff over Time\")\n    ax[1, 1].legend()\n\n    plt.tight_layout()\n    plt.show()\n\n    return all_feats\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:09.443325Z","iopub.execute_input":"2026-07-12T09:18:09.443656Z","iopub.status.idle":"2026-07-12T09:18:09.466893Z","shell.execute_reply.started":"2026-07-12T09:18:09.443628Z","shell.execute_reply":"2026-07-12T09:18:09.465914Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"feat_vec = extract_and_plot(example_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:09.468431Z","iopub.execute_input":"2026-07-12T09:18:09.468688Z","iopub.status.idle":"2026-07-12T09:18:11.273085Z","shell.execute_reply.started":"2026-07-12T09:18:09.468667Z","shell.execute_reply":"2026-07-12T09:18:11.272094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Build the Working Dataset","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(os.path.join(input_dir, 'train.csv'))\nprint(train_df.shape)\ntrain_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:11.274058Z","iopub.execute_input":"2026-07-12T09:18:11.274328Z","iopub.status.idle":"2026-07-12T09:18:11.682927Z","shell.execute_reply.started":"2026-07-12T09:18:11.274307Z","shell.execute_reply":"2026-07-12T09:18:11.682048Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Quick sanity check on a small random sample across many species.\n# This just confirms the extraction pipeline runs end-to-end -- it is\n# NOT the dataset used for modeling below (kept in separate X_sample/y_sample\n# variables so it can't silently get overwritten or reused by mistake).\nbase_path = os.path.join(input_dir, 'train_audio')\nX_sample, y_sample = [], []\n\nfor _, row in train_df.sample(20, random_state=42).iterrows():\n    species = row['ebird_code']\n    path = os.path.join(base_path, species, row['filename'])\n    if os.path.exists(path):\n        feats = safe_extract(path)\n        if feats is not None:\n            X_sample.append(feats)\n            y_sample.append(species)\n    else:\n        print(\"File not found:\", path)\n\nX_sample = np.array(X_sample)\ny_sample = np.array(y_sample)\nprint(\"Sanity-check X shape:\", X_sample.shape, \" y shape:\", y_sample.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:11.683869Z","iopub.execute_input":"2026-07-12T09:18:11.684156Z","iopub.status.idle":"2026-07-12T09:18:22.766550Z","shell.execute_reply.started":"2026-07-12T09:18:11.684127Z","shell.execute_reply":"2026-07-12T09:18:22.765600Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The full dataset has 264 species, many with only a handful of recordings each --\nnot enough per class to train and evaluate classical ML models reliably. So instead\nwe pick a small number of **well-represented** species and build a proper dataset\nfrom those.\n\n`NUM_SPECIES` below is the knob for the **\"scale to more species\"** next step: once\nyou have enough recordings per class (raise `MIN_RECORDINGS` and/or `SAMPLES_PER_SPECIES`\nif the dataset supports it), just increase this number and every downstream cell\n(feature selection, baseline models, GRU, GAN) picks it up automatically.","metadata":{}},{"cell_type":"code","source":"counts = Counter(train_df['ebird_code'])\nMIN_RECORDINGS = 80\ncommon_species = sorted([sp for sp, c in counts.items() if c >= MIN_RECORDINGS])\nprint(f\"{len(common_species)} species with >= {MIN_RECORDINGS} recordings:\")\nprint(common_species)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:22.769604Z","iopub.execute_input":"2026-07-12T09:18:22.770120Z","iopub.status.idle":"2026-07-12T09:18:22.778539Z","shell.execute_reply.started":"2026-07-12T09:18:22.770097Z","shell.execute_reply":"2026-07-12T09:18:22.777586Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Main working dataset used for everything below.\nNUM_SPECIES = 5          # <-- raise this to scale to more species\nSAMPLES_PER_SPECIES = 100\n\nselected_species = common_species[:NUM_SPECIES] if len(common_species) >= NUM_SPECIES else common_species\nprint(\"Selected species:\", selected_species)\n\nsubset_df = train_df[train_df['ebird_code'].isin(selected_species)]\n\nX, y = [], []\nfile_paths = []   # keep paths alongside features -- needed later for the GRU section\n\nfor sp in selected_species:\n    sp_df = subset_df[subset_df['ebird_code'] == sp]\n    n_samples = min(SAMPLES_PER_SPECIES, len(sp_df))\n    sp_df = sp_df.sample(n_samples, random_state=42)\n\n    for _, row in sp_df.iterrows():\n        path = os.path.join(base_path, sp, row['filename'])\n        if os.path.exists(path):\n            feats = safe_extract(path)\n            if feats is not None:\n                X.append(feats)\n                y.append(sp)\n                file_paths.append(path)\n\nX, y = np.array(X), np.array(y)\nfile_paths = np.array(file_paths)\nprint(\"X shape:\", X.shape, \" y shape:\", y.shape)\nprint(\"Counts per species:\", Counter(y))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:18:22.779371Z","iopub.execute_input":"2026-07-12T09:18:22.779600Z","iopub.status.idle":"2026-07-12T09:24:37.927687Z","shell.execute_reply.started":"2026-07-12T09:18:22.779582Z","shell.execute_reply":"2026-07-12T09:24:37.926873Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Feature Selection & Importance","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_selection import VarianceThreshold\n\nselector_var = VarianceThreshold(threshold=0.01)\nX_var = selector_var.fit_transform(X)\nprint(\"Kept features after Variance Threshold:\", X_var.shape[1], \"/\", X.shape[1])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:37.928705Z","iopub.execute_input":"2026-07-12T09:24:37.929026Z","iopub.status.idle":"2026-07-12T09:24:37.938894Z","shell.execute_reply.started":"2026-07-12T09:24:37.929002Z","shell.execute_reply":"2026-07-12T09:24:37.938062Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.feature_selection import SelectKBest, f_classif\n\nselector_anova = SelectKBest(score_func=f_classif, k=10)\nX_best = selector_anova.fit_transform(X, y)\nanova_scores = selector_anova.scores_\nprint(\"Top 10 features selected via ANOVA F-test:\", X_best.shape[1])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:37.939969Z","iopub.execute_input":"2026-07-12T09:24:37.940363Z","iopub.status.idle":"2026-07-12T09:24:37.958279Z","shell.execute_reply.started":"2026-07-12T09:24:37.940339Z","shell.execute_reply":"2026-07-12T09:24:37.957385Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nrf_importance_model = RandomForestClassifier(n_estimators=200, random_state=42, n_jobs=-1)\nrf_importance_model.fit(X, y)\n\nrf_importances = rf_importance_model.feature_importances_\nrf_rank = np.argsort(rf_importances)[::-1]\n\n# Fixed: the original notebook named this \"top10_idx\" but sliced [:20],\n# then later fed it into a feature set literally called \"Top10\". Now the\n# name matches the content.\ntop10_idx = rf_rank[:10]\n\nprint(\"Top 10 features by Random Forest importance:\")\nfor rank, idx in enumerate(top10_idx, start=1):\n    print(f\"{rank}. {FEATURE_NAMES[idx]} (importance={rf_importances[idx]:.4f})\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:37.959185Z","iopub.execute_input":"2026-07-12T09:24:37.959435Z","iopub.status.idle":"2026-07-12T09:24:38.545629Z","shell.execute_reply.started":"2026-07-12T09:24:37.959416Z","shell.execute_reply":"2026-07-12T09:24:38.544914Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df = pd.DataFrame({\n    \"Feature_Index\": np.arange(len(FEATURE_NAMES)),\n    \"Feature_Name\": FEATURE_NAMES,\n    \"ANOVA_Fscore\": anova_scores,\n    \"RF_Importance\": rf_importances\n})\n\ndf[\"ANOVA_Fscore_norm\"] = df[\"ANOVA_Fscore\"] / df[\"ANOVA_Fscore\"].max()\ndf[\"RF_Importance_norm\"] = df[\"RF_Importance\"] / df[\"RF_Importance\"].max()\n\ntop_anova = set(df.nlargest(10, \"ANOVA_Fscore\").Feature_Index)\ntop_rf = set(df.nlargest(10, \"RF_Importance\").Feature_Index)\n\n# Fixed: previously hardcoded as [1, 5, 6, 7] based on a one-off manual read\n# of this table. Now computed directly, so it stays correct if the data changes.\noverlap_idx = sorted(top_anova & top_rf)\n\ndf[\"Selected_by_Both\"] = df[\"Feature_Index\"].apply(lambda i: \"✅\" if i in overlap_idx else \"\")\n\ndf_sorted = df.sort_values([\"Selected_by_Both\", \"ANOVA_Fscore\"], ascending=[False, False])\nprint(\"Features selected by both ANOVA and RF top-10:\", overlap_idx)\ndf_sorted.head(20)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:38.546555Z","iopub.execute_input":"2026-07-12T09:24:38.546800Z","iopub.status.idle":"2026-07-12T09:24:38.577681Z","shell.execute_reply.started":"2026-07-12T09:24:38.546774Z","shell.execute_reply":"2026-07-12T09:24:38.576921Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_chart = df.sort_values(\"RF_Importance\", ascending=False).reset_index(drop=True)\n\ntopN = 20\nplt.figure(figsize=(14, 7))\nbar_width = 0.4\nidx_range = np.arange(topN)\n\nplt.bar(idx_range, df_chart[\"ANOVA_Fscore_norm\"].head(topN), bar_width,\n        label=\"ANOVA F-score (normalized)\", alpha=0.7, color=\"skyblue\")\nplt.bar(idx_range + bar_width, df_chart[\"RF_Importance_norm\"].head(topN), bar_width,\n        label=\"RF Importance (normalized)\", alpha=0.7, color=\"orange\")\n\nfor i in range(topN):\n    if df_chart.loc[i, \"Selected_by_Both\"] == \"✅\":\n        y_pos = max(df_chart[\"ANOVA_Fscore_norm\"].iloc[i], df_chart[\"RF_Importance_norm\"].iloc[i]) + 0.05\n        plt.text(idx_range[i] + bar_width / 2, y_pos, \"⭐\", ha=\"center\", fontsize=14, color=\"red\")\n\nplt.xticks(idx_range + bar_width / 2, df_chart[\"Feature_Name\"].head(topN), rotation=45, ha=\"right\")\nplt.ylabel(\"Normalized Score\")\nplt.title(\"Top 20 Features: ANOVA vs Random Forest (⭐ = Selected by Both Top-10 Lists)\")\nplt.legend()\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:38.578833Z","iopub.execute_input":"2026-07-12T09:24:38.579185Z","iopub.status.idle":"2026-07-12T09:24:38.990538Z","shell.execute_reply.started":"2026-07-12T09:24:38.579156Z","shell.execute_reply":"2026-07-12T09:24:38.989678Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Baseline Model Comparison (All Features)\n\nReporting both **accuracy** and **macro-F1** from here on, since macro-F1 treats every species equally regardless of how many recordings it has -- a fairer view than accuracy alone when classes are imbalanced.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.svm import SVC\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import classification_report, accuracy_score, f1_score\n\ndef report_scores(name, y_true, y_pred, verbose=True):\n    acc = accuracy_score(y_true, y_pred)\n    macro_f1 = f1_score(y_true, y_pred, average=\"macro\")\n    print(f\"{name} | Accuracy: {acc:.3f} | Macro-F1: {macro_f1:.3f}\")\n    if verbose:\n        print(classification_report(y_true, y_pred))\n    return acc, macro_f1\n\nX_train, X_test, y_train, y_test = train_test_split(\n    X, y, test_size=0.2, stratify=y, random_state=42\n)\n\nbaseline_models = {\n    \"RandomForest\": RandomForestClassifier(n_estimators=200, random_state=42),\n    \"SVM (RBF)\": SVC(kernel=\"rbf\", probability=True, random_state=42),\n    \"KNN\": KNeighborsClassifier(n_neighbors=5),\n    \"LogReg\": LogisticRegression(max_iter=2000, random_state=42)\n}\n\nbaseline_results = {}\nfor name, clf in baseline_models.items():\n    clf.fit(X_train, y_train)\n    y_pred = clf.predict(X_test)\n    print()\n    baseline_results[name] = report_scores(name, y_test, y_pred)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:38.991509Z","iopub.execute_input":"2026-07-12T09:24:38.991811Z","iopub.status.idle":"2026-07-12T09:24:50.611773Z","shell.execute_reply.started":"2026-07-12T09:24:38.991788Z","shell.execute_reply":"2026-07-12T09:24:50.611121Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Comparing Feature Subsets\n\nDoes the full 37-feature vector actually help, or would a smaller, cheaper feature set do just as well?","metadata":{}},{"cell_type":"code","source":"X_full = X\nX_top10 = X[:, top10_idx]\nX_overlap = X[:, overlap_idx] if len(overlap_idx) > 0 else X[:, top10_idx[:4]]\n\nfeature_sets = {\n    \"Full (37)\": X_full,\n    \"Top10\": X_top10,\n    f\"Overlap ({X_overlap.shape[1]})\": X_overlap\n}\n\nidx_all = np.arange(len(y))\ntrain_idx, test_idx, y_train, y_test = train_test_split(\n    idx_all, y, test_size=0.2, stratify=y, random_state=42\n)\n\nsubset_models = {\n    \"RandomForest\": RandomForestClassifier(n_estimators=200, random_state=42),\n    \"LogReg\": LogisticRegression(max_iter=2000, random_state=42),\n    \"KNN\": KNeighborsClassifier(n_neighbors=5),\n    \"SVM(RBF)\": SVC(kernel=\"rbf\", gamma=\"scale\", random_state=42)\n}\n\nresults_holdout = {}\n\nfor fs_name, X_set in feature_sets.items():\n    X_tr, X_te = X_set[train_idx], X_set[test_idx]\n    print(f\"\\n=== Feature set: {fs_name} ===\")\n    results_holdout[fs_name] = {}\n    for model_name, model in subset_models.items():\n        model.fit(X_tr, y_train)\n        y_pred = model.predict(X_te)\n        acc, macro_f1 = report_scores(model_name, y_test, y_pred, verbose=False)\n        results_holdout[fs_name][model_name] = (acc, macro_f1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:24:50.612481Z","iopub.execute_input":"2026-07-12T09:24:50.612738Z","iopub.status.idle":"2026-07-12T09:25:04.605766Z","shell.execute_reply.started":"2026-07-12T09:24:50.612715Z","shell.execute_reply":"2026-07-12T09:25:04.605074Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6b. Hyperparameter Tuning (GridSearchCV / RandomizedSearchCV)\n\nSections 5-6 used \"reasonable defaults\" for every model (`n_estimators=200`,\n`n_neighbors=5`, default `C`, etc.). Here we actually search for better\nsettings, using the same held-out `X_train`/`y_train`/`X_test`/`y_test` split\nfrom Section 5 (Full 37-feature set) so results are directly comparable to\n`baseline_results`.\n\nTwo search strategies, matched to each model's parameter space:\n- **`GridSearchCV`** for **LogReg** and **KNN** -- their useful hyperparameters\n  are a handful of discrete choices (a `C` value, a distance metric, a\n  weighting scheme), small enough to exhaustively enumerate.\n- **`RandomizedSearchCV`** for **RandomForest** and **SVM(RBF)** -- their\n  hyperparameters interact and span wide/continuous ranges (`C`, `gamma`,\n  tree depth and split thresholds), where exhaustive grid search would need\n  far more combinations than are worth evaluating on a dataset this size;\n  random sampling covers the space with a fixed, controllable budget instead.\n\nEvery search optimizes **macro-F1** via 5-fold stratified CV (consistent\nwith how every other model comparison in this notebook is scored), and\nLogReg/KNN/SVM are wrapped in a `StandardScaler` pipeline since (unlike\nRandomForest) they're scale-sensitive.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV, RandomizedSearchCV, StratifiedKFold\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import StandardScaler\nfrom scipy.stats import randint, loguniform\n\ntune_cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\ntuning_searches = {}\n\n# --- RandomForest: RandomizedSearchCV -- wide, mixed discrete/continuous space ---\nrf_search = RandomizedSearchCV(\n    RandomForestClassifier(random_state=42),\n    param_distributions={\n        \"n_estimators\": randint(100, 400),\n        \"max_depth\": [None, 5, 10, 20],\n        \"min_samples_split\": randint(2, 10),\n        \"min_samples_leaf\": randint(1, 5),\n        \"max_features\": [\"sqrt\", \"log2\", None],\n    },\n    n_iter=25, scoring=\"f1_macro\", cv=tune_cv, random_state=42, n_jobs=-1,\n)\nrf_search.fit(X_train, y_train)\ntuning_searches[\"RandomForest\"] = rf_search\n\n# --- LogReg: GridSearchCV -- one discrete axis (regularization strength) ---\nlogreg_pipe = Pipeline([\n    (\"scaler\", StandardScaler()),\n    (\"clf\", LogisticRegression(max_iter=2000, random_state=42)),\n])\nlogreg_search = GridSearchCV(\n    logreg_pipe,\n    param_grid={\"clf__C\": [0.01, 0.1, 1, 10, 100]},\n    scoring=\"f1_macro\", cv=tune_cv, n_jobs=-1,\n)\nlogreg_search.fit(X_train, y_train)\ntuning_searches[\"LogReg\"] = logreg_search\n\n# --- KNN: GridSearchCV -- small, fully discrete grid ---\nknn_pipe = Pipeline([\n    (\"scaler\", StandardScaler()),\n    (\"clf\", KNeighborsClassifier()),\n])\nknn_search = GridSearchCV(\n    knn_pipe,\n    param_grid={\n        \"clf__n_neighbors\": [3, 5, 7, 9, 11],\n        \"clf__weights\": [\"uniform\", \"distance\"],\n        \"clf__metric\": [\"euclidean\", \"manhattan\"],\n    },\n    scoring=\"f1_macro\", cv=tune_cv, n_jobs=-1,\n)\nknn_search.fit(X_train, y_train)\ntuning_searches[\"KNN\"] = knn_search\n\n# --- SVM(RBF): RandomizedSearchCV -- continuous C/gamma, log-scaled ---\nsvm_pipe = Pipeline([\n    (\"scaler\", StandardScaler()),\n    (\"clf\", SVC(kernel=\"rbf\", random_state=42)),\n])\nsvm_search = RandomizedSearchCV(\n    svm_pipe,\n    param_distributions={\n        \"clf__C\": loguniform(1e-1, 1e3),\n        \"clf__gamma\": loguniform(1e-4, 1e0),\n    },\n    n_iter=25, scoring=\"f1_macro\", cv=tune_cv, random_state=42, n_jobs=-1,\n)\nsvm_search.fit(X_train, y_train)\ntuning_searches[\"SVM (RBF)\"] = svm_search\n\nprint(\"=== Best CV macro-F1 and params found for each model ===\")\nfor name, search in tuning_searches.items():\n    print(f\"{name:14s} | best CV Macro-F1: {search.best_score_:.3f} | best params: {search.best_params_}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:25:04.606495Z","iopub.execute_input":"2026-07-12T09:25:04.606755Z","iopub.status.idle":"2026-07-12T09:25:56.917868Z","shell.execute_reply.started":"2026-07-12T09:25:04.606729Z","shell.execute_reply":"2026-07-12T09:25:56.916992Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Evaluate the tuned models on the same held-out test set baseline_results\n# used, so \"baseline vs. tuned\" is an apples-to-apples comparison.\nprint(\"=== Baseline (defaults) vs. Tuned (Grid/RandomizedSearchCV), held-out test set ===\")\ntuned_results = {}\nfor name, search in tuning_searches.items():\n    y_pred = search.best_estimator_.predict(X_test)\n    acc, macro_f1 = report_scores(f\"{name} (tuned)\", y_test, y_pred, verbose=False)\n    tuned_results[name] = (acc, macro_f1)\n\nprint()\nfor name in tuning_searches:\n    base_acc, base_f1 = baseline_results[name]\n    tuned_acc, tuned_f1 = tuned_results[name]\n    delta_acc = tuned_acc - base_acc\n    print(\n        f\"{name:14s} | Baseline Acc/F1: {base_acc:.3f}/{base_f1:.3f} \"\n        f\"| Tuned Acc/F1: {tuned_acc:.3f}/{tuned_f1:.3f} | Acc change: {delta_acc:+.3f}\"\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:25:56.918817Z","iopub.execute_input":"2026-07-12T09:25:56.919325Z","iopub.status.idle":"2026-07-12T09:25:56.975865Z","shell.execute_reply.started":"2026-07-12T09:25:56.919302Z","shell.execute_reply":"2026-07-12T09:25:56.975132Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Class Balancing with SMOTE + Cross-Validation\n\n**Note on the data going into this section:** the working dataset built in\nSection 3 uses a fixed `SAMPLES_PER_SPECIES = 100`, so it comes in **already\nperfectly balanced** (100/100/100/100/100) whenever enough recordings exist\nfor every selected species. If SMOTE (or the cGAN below) runs on data that's\nalready balanced, there's no minority class left to correct -- both methods\ncompute that zero synthetic samples are needed and just hand back the\noriginal data unchanged. Comparing \"SMOTE vs. cGAN\" on that would be\ncomparing a method against itself, not against another method.\n\nSo for *this section only*, we build a synthetically **imbalanced** copy of\nthe dataset (`X_imb`, `y_imb`) by randomly downsampling two of the five\nspecies down to a small count, keeping the rest at full size. That gives\nSMOTE and, later, the cGAN an actual imbalance to fix, so their comparison\nmeans something. Everything upstream (feature selection, Section 5/6\nbaselines) still uses the original balanced `X`/`y` -- only this\nclass-balancing comparison uses the artificially imbalanced version.\n\n5-fold stratified CV gives a more robust estimate than a single train/test\nsplit; we score on both accuracy and macro-F1.","metadata":{}},{"cell_type":"code","source":"# Build a synthetically imbalanced copy of the dataset for this comparison\n# only -- see the note above. Two of the five species are downsampled to a\n# small minority count; the rest keep their full sample count.\nrng_imb = np.random.RandomState(42)\nMINORITY_SIZE = 15\nminority_species = selected_species[:2]   # first two species become the \"rare\" classes\n\nimb_idx = []\nfor cls in np.unique(y):\n    cls_idx = np.where(y == cls)[0]\n    if cls in minority_species:\n        cls_idx = rng_imb.choice(cls_idx, size=min(MINORITY_SIZE, len(cls_idx)), replace=False)\n    imb_idx.append(cls_idx)\nimb_idx = np.concatenate(imb_idx)\nrng_imb.shuffle(imb_idx)\n\nX_imb, y_imb = X[imb_idx], y[imb_idx]\nprint(\"Imbalanced dataset for SMOTE/cGAN comparison:\", X_imb.shape, \"Counts:\", Counter(y_imb))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:25:56.976762Z","iopub.execute_input":"2026-07-12T09:25:56.977124Z","iopub.status.idle":"2026-07-12T09:25:56.985309Z","shell.execute_reply.started":"2026-07-12T09:25:56.977102Z","shell.execute_reply":"2026-07-12T09:25:56.984462Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold, cross_validate\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.pipeline import make_pipeline\nfrom imblearn.over_sampling import SMOTE\n\n# Guard k_neighbors against very small classes (SMOTE requires\n# k_neighbors < smallest class size, otherwise it raises an error).\nmin_class_size = min(Counter(y_imb).values())\nsmote = SMOTE(random_state=42, k_neighbors=min(5, max(1, min_class_size - 1)))\nX_bal, y_bal = smote.fit_resample(X_imb, y_imb)\nprint(\"Balanced dataset (SMOTE):\", X_bal.shape, \"Counts:\", Counter(y_bal))\n\nfeature_sets_bal = {\n    \"Full (37)\": X_bal,\n    \"Top10\": X_bal[:, top10_idx],\n    f\"Overlap ({len(overlap_idx) if overlap_idx else 4})\":\n        X_bal[:, overlap_idx] if len(overlap_idx) > 0 else X_bal[:, top10_idx[:4]]\n}\n\ncv_models = {\n    \"RandomForest\": RandomForestClassifier(n_estimators=200, random_state=42),\n    \"LogReg\": make_pipeline(StandardScaler(), LogisticRegression(max_iter=2000, random_state=42)),\n    \"KNN\": make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=5)),\n    \"SVM(RBF)\": make_pipeline(StandardScaler(), SVC(kernel=\"rbf\", gamma=\"scale\", random_state=42))\n}\n\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\nresults_cv_smote = {}\n\nfor fs_name, X_set in feature_sets_bal.items():\n    print(f\"\\n=== Feature Set: {fs_name} ===\")\n    results_cv_smote[fs_name] = {}\n    for model_name, model in cv_models.items():\n        scores = cross_validate(model, X_set, y_bal, cv=skf,\n                                 scoring=[\"accuracy\", \"f1_macro\"], n_jobs=-1)\n        acc_mean, acc_std = scores[\"test_accuracy\"].mean(), scores[\"test_accuracy\"].std()\n        f1_mean, f1_std = scores[\"test_f1_macro\"].mean(), scores[\"test_f1_macro\"].std()\n        results_cv_smote[fs_name][model_name] = (acc_mean, f1_mean)\n        print(f\"{model_name} | Acc: {acc_mean:.3f} (+/- {acc_std:.3f}) | \"\n              f\"Macro-F1: {f1_mean:.3f} (+/- {f1_std:.3f})\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:25:56.986218Z","iopub.execute_input":"2026-07-12T09:25:56.987578Z","iopub.status.idle":"2026-07-12T09:26:01.560036Z","shell.execute_reply.started":"2026-07-12T09:25:56.987552Z","shell.execute_reply":"2026-07-12T09:26:01.559123Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Deep Learning Alternative: GRU on Mel-Spectrogram Sequences\n\nThe sections above rely on **handcrafted, time-averaged** features (a single 37-dim\nvector per clip). That throws away how the sound evolves over time. A **GRU**\n(Gated Recurrent Unit) is a natural fit here: it reads a clip as a *sequence* of\nmel-spectrogram frames (time steps) and can learn temporal patterns directly,\nwithout hand-designing them.\n\nThis works well for this dataset because:\n- Clips are naturally sequential (time x frequency), which is exactly what a GRU expects.\n- The species subset here is small (5 classes, a few hundred clips) -- a lightweight\n  GRU trains quickly even on CPU, unlike a full CNN on raw spectrogram images.\n\nEach clip is converted to a fixed-length sequence of mel-spectrogram frames\n(padded/truncated to `MAX_FRAMES` time steps), and the GRU classifies the sequence.\n\n**Note:** the first run of this model showed a large train/val accuracy gap (train climbing past 0.8 while val stayed in the 0.5-0.65 range) -- classic overfitting on a few hundred clips. The training cell below adds `recurrent_dropout`, L2 weight decay, `EarlyStopping`, and `ReduceLROnPlateau` to close that gap; see the comment in that cell for details. `recurrent_dropout > 0` does disable the fast cuDNN GRU implementation, so this trains a bit slower per epoch than the original -- a fine trade for a small dataset where overfitting is the bigger risk.","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models\nfrom sklearn.preprocessing import LabelEncoder\n\ntf.random.set_seed(42)\n\nN_MELS = 64\nMAX_FRAMES = 130   # ~4 seconds of audio at sr=16000, hop=512\n\ndef extract_melspec_sequence(path, sr=16000, n_mels=N_MELS, hop=512, max_frames=MAX_FRAMES):\n    \"\"\"Returns a (max_frames, n_mels) sequence: mel-spectrogram frames over time,\n    padded with zeros or truncated to a fixed length so all clips have the same shape.\"\"\"\n    y_audio, _ = librosa.load(path, sr=sr)\n    mel = librosa.feature.melspectrogram(y=y_audio, sr=sr, n_mels=n_mels, hop_length=hop)\n    mel_db = librosa.power_to_db(mel, ref=np.max).T   # (time, n_mels)\n\n    if mel_db.shape[0] >= max_frames:\n        mel_db = mel_db[:max_frames]\n    else:\n        pad_width = max_frames - mel_db.shape[0]\n        mel_db = np.pad(mel_db, ((0, pad_width), (0, 0)), mode=\"constant\", constant_values=mel_db.min())\n    return mel_db\n\nprint(\"Building mel-spectrogram sequences for the GRU (this can take a few minutes)...\")\n# file_paths already passed safe_extract() once (in the dataset-building cell),\n# but wrap this pass too -- a different feature path can occasionally hit a\n# decoder edge case the first pass didn't, and we'd rather drop one clip than\n# crash np.stack() over a shape mismatch.\nseq_list, y_seq_list = [], []\nfor p, label in zip(file_paths, y):\n    try:\n        seq_list.append(extract_melspec_sequence(p))\n        y_seq_list.append(label)\n    except Exception as e:\n        print(f\"Skipping for GRU (failed to decode): {p} -- {type(e).__name__}: {e}\")\n\nX_seq = np.stack(seq_list)\ny_gru_labels = np.array(y_seq_list)\nprint(\"X_seq shape:\", X_seq.shape, \"-> (n_samples, time_steps, n_mels)\")\n\n# Per-frequency-bin normalization helps GRU training converge faster\nseq_mean = X_seq.mean(axis=(0, 1), keepdims=True)\nseq_std = X_seq.std(axis=(0, 1), keepdims=True) + 1e-6\nX_seq_norm = (X_seq - seq_mean) / seq_std\n\nlabel_enc = LabelEncoder()\ny_int = label_enc.fit_transform(y_gru_labels)\nnum_classes = len(label_enc.classes_)\n\nXs_train, Xs_test, ys_train, ys_test = train_test_split(\n    X_seq_norm, y_int, test_size=0.2, stratify=y_int, random_state=42\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:26:01.561106Z","iopub.execute_input":"2026-07-12T09:26:01.561444Z","iopub.status.idle":"2026-07-12T09:27:35.163118Z","shell.execute_reply.started":"2026-07-12T09:26:01.561415Z","shell.execute_reply":"2026-07-12T09:27:35.162102Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---- Regularization to close the train/val accuracy gap seen above ----\n# The original GRU (64 -> 32 units, single Dropout(0.3), fit for a fixed 30\n# epochs) memorizes the ~340 training clips well before validation catches\n# up -- train accuracy climbs past 0.8 while val bounces around 0.5-0.65.\n# With only a few hundred clips and 5 classes, that gap is classic\n# overfitting, not a hard problem. Four changes, applied together:\n#   1. recurrent_dropout on both GRU layers -- dropout applied to the\n#      recurrent (timestep-to-timestep) connections, not just the output,\n#      which regularizes the sequence model itself rather than just the\n#      dense head.\n#   2. L2 weight decay on the GRU and Dense kernels -- discourages any\n#      single unit from growing large weights to memorize a few clips.\n#   3. EarlyStopping on val_loss with restore_best_weights -- stop at the\n#      epoch where validation actually stopped improving instead of\n#      training a fixed 30 epochs regardless, and roll back to that point.\n#   4. ReduceLROnPlateau -- shrink the learning rate when val_loss plateaus,\n#      so the model fine-tunes instead of overshooting once it's close.\nfrom tensorflow.keras import regularizers\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau\n\nL2 = 1e-4\n\ndef build_gru_model(time_steps, n_mels, num_classes):\n    model = models.Sequential([\n        layers.Input(shape=(time_steps, n_mels)),\n        layers.Masking(mask_value=0.0),\n        layers.GRU(64, return_sequences=True,\n                    dropout=0.3, recurrent_dropout=0.2,\n                    kernel_regularizer=regularizers.l2(L2)),\n        layers.GRU(32,\n                    dropout=0.3, recurrent_dropout=0.2,\n                    kernel_regularizer=regularizers.l2(L2)),\n        layers.Dense(32, activation=\"relu\", kernel_regularizer=regularizers.l2(L2)),\n        layers.Dropout(0.4),\n        layers.Dense(num_classes, activation=\"softmax\")\n    ])\n    model.compile(optimizer=\"adam\", loss=\"sparse_categorical_crossentropy\", metrics=[\"accuracy\"])\n    return model\n\ngru_model = build_gru_model(MAX_FRAMES, N_MELS, num_classes)\ngru_model.summary()\n\ncallbacks = [\n    EarlyStopping(monitor=\"val_loss\", patience=8, restore_best_weights=True),\n    ReduceLROnPlateau(monitor=\"val_loss\", factor=0.5, patience=4, min_lr=1e-5),\n]\n\n# Set a generous epoch ceiling -- EarlyStopping decides the real stopping\n# point and restores the best-val-loss weights, so this isn't \"training\n# longer,\" it's \"stop training at the right time instead of a fixed 30.\"\nhistory = gru_model.fit(\n    Xs_train, ys_train,\n    validation_split=0.15,\n    epochs=100,\n    batch_size=16,\n    callbacks=callbacks,\n    verbose=1\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:27:35.164215Z","iopub.execute_input":"2026-07-12T09:27:35.164610Z","iopub.status.idle":"2026-07-12T09:30:49.399882Z","shell.execute_reply.started":"2026-07-12T09:27:35.164582Z","shell.execute_reply":"2026-07-12T09:30:49.399071Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_pred_gru_int = np.argmax(gru_model.predict(Xs_test), axis=1)\ny_pred_gru = label_enc.inverse_transform(y_pred_gru_int)\ny_test_gru = label_enc.inverse_transform(ys_test)\n\ngru_acc, gru_f1 = report_scores(\"GRU (mel-spectrogram sequence)\", y_test_gru, y_pred_gru)\n\nplt.figure(figsize=(8, 4))\nplt.plot(history.history[\"accuracy\"], label=\"Train Accuracy\")\nplt.plot(history.history[\"val_accuracy\"], label=\"Val Accuracy\")\nplt.title(\"GRU Training Accuracy\")\nplt.xlabel(\"Epoch\")\nplt.ylabel(\"Accuracy\")\nplt.legend()\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:30:49.401521Z","iopub.execute_input":"2026-07-12T09:30:49.401842Z","iopub.status.idle":"2026-07-12T09:30:52.383468Z","shell.execute_reply.started":"2026-07-12T09:30:49.401813Z","shell.execute_reply":"2026-07-12T09:30:52.382565Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8b. Deep Learning Alternative: True 2D CNN on Mel-Spectrogram Images\n\nThe GRU above treats each clip as a *sequence* of frames (1D over time, with\nfrequency as the feature dimension at each step). A 2D CNN instead treats the\nmel-spectrogram as an **image** (time x frequency) and slides small 2D filters\nacross both axes at once. This lets it pick up **local time-frequency patterns**\ndirectly -- a short trill, a sharp harmonic stack, a frequency sweep -- the kind\nof localized texture a GRU (which only sees one frequency slice per time step)\ncan smear across time steps or miss entirely.\n\nWe reuse the exact same `X_seq_norm` tensor already built for the GRU\n(`(n_samples, MAX_FRAMES, N_MELS)`, per-bin normalized) -- no need to\nre-extract audio. For a CNN we just add a channel dimension, treating it as a\nsingle-channel `(time, frequency, 1)` image, and swap the GRU layers for\n`Conv2D` + `MaxPooling2D` blocks.\n\nTo help the CNN generalize on this small sample (a few hundred clips), we apply mild **SpecAugment**-style augmentation to the *training* split only: a random block of time steps and a random block of frequency bins are zeroed out on each pass, freshly re-rolled every epoch via a `tf.data` pipeline. This is cheap, doesn't touch the underlying audio, and is a standard regularizer for spectrogram CNNs -- unlike SMOTE/GAN (which synthesize new feature vectors), it works directly on the image input and only masks a small fraction of each clip, so labels stay unambiguous.","metadata":{}},{"cell_type":"code","source":"# Reuse the GRU's mel-spectrogram tensor as a single-channel \"image\":\n# (n_samples, time_steps, n_mels) -> (n_samples, time_steps, n_mels, 1)\nX_img_train_full = Xs_train[..., np.newaxis]\nX_img_test = Xs_test[..., np.newaxis]\nprint(\"CNN input shape:\", X_img_train_full.shape, \"-> (n_samples, time_steps, n_mels, channels)\")\n\n# ---- SpecAugment-style augmentation (mild) ----\n# Randomly zero out one short block of time steps and one short block of\n# frequency bins per sample. Re-rolled fresh every epoch via tf.data (not\n# baked into the arrays once), so the CNN sees a slightly different masked\n# view of each clip on every pass. Values are already per-bin normalized\n# (mean ~0), so zeroing approximates SpecAugment's usual \"mask with the mean\"\n# convention rather than creating an artificial hard-black region.\nMAX_TIME_MASK = 15    # up to ~15 frames (~0.5s at sr=16000, hop=512) zeroed\nMAX_FREQ_MASK = 8     # up to ~8 mel bins zeroed\nN_TIME_MASKS = 1\nN_FREQ_MASKS = 1\n\ndef spec_augment(spec, label):\n    \"\"\"spec: (time_steps, n_mels, 1) tensor. Only ever applied to the\n    training split -- validation/test stay unaugmented so evaluation reflects\n    real clips.\"\"\"\n    time_steps = tf.shape(spec)[0]\n    n_mels = tf.shape(spec)[1]\n\n    for _ in range(N_TIME_MASKS):\n        t = tf.random.uniform([], 0, MAX_TIME_MASK + 1, dtype=tf.int32)\n        t0 = tf.random.uniform([], 0, tf.maximum(time_steps - t, 1), dtype=tf.int32)\n        time_mask = tf.concat([\n            tf.ones([t0, n_mels, 1]),\n            tf.zeros([t, n_mels, 1]),\n            tf.ones([time_steps - t0 - t, n_mels, 1])\n        ], axis=0)\n        spec = spec * time_mask\n\n    for _ in range(N_FREQ_MASKS):\n        f = tf.random.uniform([], 0, MAX_FREQ_MASK + 1, dtype=tf.int32)\n        f0 = tf.random.uniform([], 0, tf.maximum(n_mels - f, 1), dtype=tf.int32)\n        freq_mask = tf.concat([\n            tf.ones([time_steps, f0, 1]),\n            tf.zeros([time_steps, f, 1]),\n            tf.ones([time_steps, n_mels - f0 - f, 1])\n        ], axis=1)\n        spec = spec * freq_mask\n\n    return spec, label\n\n# Carve out the same ~15% validation slice model.fit(validation_split=...) used\n# for the GRU, but explicitly -- tf.data pipelines don't take validation_split.\nX_cnn_train, X_cnn_val, ys_cnn_train, ys_cnn_val = train_test_split(\n    X_img_train_full, ys_train, test_size=0.15, stratify=ys_train, random_state=42\n)\n\nAUTOTUNE = tf.data.AUTOTUNE\nBATCH_SIZE_CNN = 16\n\ntrain_ds = (\n    tf.data.Dataset.from_tensor_slices((X_cnn_train, ys_cnn_train))\n    .shuffle(len(X_cnn_train), seed=42, reshuffle_each_iteration=True)\n    .map(spec_augment, num_parallel_calls=AUTOTUNE)   # augmentation: train only\n    .batch(BATCH_SIZE_CNN)\n    .prefetch(AUTOTUNE)\n)\nval_ds = (\n    tf.data.Dataset.from_tensor_slices((X_cnn_val, ys_cnn_val))   # no augmentation\n    .batch(BATCH_SIZE_CNN)\n    .prefetch(AUTOTUNE)\n)\n\ndef build_cnn_model(time_steps, n_mels, num_classes):\n    model = models.Sequential([\n        layers.Input(shape=(time_steps, n_mels, 1)),\n        layers.Conv2D(16, (3, 3), activation=\"relu\", padding=\"same\"),\n        layers.MaxPooling2D((2, 2)),\n        layers.Conv2D(32, (3, 3), activation=\"relu\", padding=\"same\"),\n        layers.MaxPooling2D((2, 2)),\n        layers.Conv2D(64, (3, 3), activation=\"relu\", padding=\"same\"),\n        layers.GlobalAveragePooling2D(),\n        layers.Dense(32, activation=\"relu\"),\n        layers.Dropout(0.3),\n        layers.Dense(num_classes, activation=\"softmax\")\n    ])\n    model.compile(optimizer=\"adam\", loss=\"sparse_categorical_crossentropy\", metrics=[\"accuracy\"])\n    return model\n\ncnn_model = build_cnn_model(MAX_FRAMES, N_MELS, num_classes)\ncnn_model.summary()\n\ncnn_history = cnn_model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=30,\n    verbose=1\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:30:52.384462Z","iopub.execute_input":"2026-07-12T09:30:52.384809Z","iopub.status.idle":"2026-07-12T09:31:36.903429Z","shell.execute_reply.started":"2026-07-12T09:30:52.384779Z","shell.execute_reply":"2026-07-12T09:31:36.902655Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_pred_cnn_int = np.argmax(cnn_model.predict(X_img_test), axis=1)\ny_pred_cnn = label_enc.inverse_transform(y_pred_cnn_int)\ny_test_cnn = label_enc.inverse_transform(ys_test)  # same split/order as ys_test used for GRU\n\ncnn_acc, cnn_f1 = report_scores(\"CNN (2D conv on mel-spectrogram image)\", y_test_cnn, y_pred_cnn)\n\nplt.figure(figsize=(8, 4))\nplt.plot(cnn_history.history[\"accuracy\"], label=\"Train Accuracy\")\nplt.plot(cnn_history.history[\"val_accuracy\"], label=\"Val Accuracy\")\nplt.title(\"2D CNN Training Accuracy\")\nplt.xlabel(\"Epoch\")\nplt.ylabel(\"Accuracy\")\nplt.legend()\nplt.tight_layout()\nplt.show()\n\nprint(f\"\\n=== GRU vs 2D CNN (same train/test split, {num_classes} classes) ===\")\nprint(f\"GRU | Acc: {gru_acc:.3f} | Macro-F1: {gru_f1:.3f}\")\nprint(f\"CNN | Acc: {cnn_acc:.3f} | Macro-F1: {cnn_f1:.3f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:31:36.904905Z","iopub.execute_input":"2026-07-12T09:31:36.905233Z","iopub.status.idle":"2026-07-12T09:31:37.521204Z","shell.execute_reply.started":"2026-07-12T09:31:36.905205Z","shell.execute_reply":"2026-07-12T09:31:37.520274Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8c. Transfer Learning: YAMNet Embeddings (frame-level, not from-scratch)\n\nThe GRU and CNN above landed at **~0.48-0.49** accuracy on 5 classes (chance =\n0.20) -- only marginally ahead of the plain **Logistic Regression on 37 hand\nfeatures (0.566)** from Section 5. That's a real, informative result, not just\na disappointing one: **~400 training clips is not enough for a randomly\ninitialized deep net to learn its own useful time-frequency filters.**\nGRUs/CNNs trained from scratch typically need thousands of examples per class\nbefore they beat hand-engineered features + a linear model -- and that's\nexactly the pattern in the numbers above.\n\nThe standard fix used across essentially every strong Cornell Birdcall /\nBirdCLEF solution is **transfer learning**: don't train filters from\nscratch -- start from a network already pretrained on a large, general-purpose\naudio corpus, and train only a small classifier head on top. We use\n**YAMNet** (Google, pretrained on **AudioSet**: 2M+ 10-second clips across 521\nsound classes, including many bird/animal categories), loaded frozen from\nTensorFlow Hub.\n\nThis also fixes a second, independent problem for free: YAMNet slices each\nclip into **~0.96s frames** (hop 0.48s) and returns one 1024-dim embedding\nper frame. Instead of averaging those down to a single vector per clip (which\nis what the handcrafted-feature pipeline does), we keep **every frame as its\nown training example**, sharing its parent clip's label. That turns ~400\nclips into several thousand frame-level examples with zero extra audio\ncollected -- effectively free data augmentation. At test time we vote a\nclip's frame-level predictions back up to a single clip-level label, so the\nresult stays directly comparable to the GRU/CNN accuracy and macro-F1 above.\n\nNote: because clips are split (not frames), no clip's frames appear on both\nsides of train/test -- otherwise near-duplicate frames from the same\nrecording would leak across the split and inflate the score.","metadata":{}},{"cell_type":"code","source":"# YAMNet ships via tensorflow_hub, not tensorflow itself.\n!pip install -q tensorflow_hub\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:31:37.525372Z","iopub.execute_input":"2026-07-12T09:31:37.525731Z","iopub.status.idle":"2026-07-12T09:31:41.445871Z","shell.execute_reply.started":"2026-07-12T09:31:37.525707Z","shell.execute_reply":"2026-07-12T09:31:41.444520Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow_hub as hub\n\nprint(\"Loading pretrained YAMNet (AudioSet) from TensorFlow Hub -- one-time download...\")\nyamnet_model = hub.load(\"https://tfhub.dev/google/yamnet/1\")\n\nYAMNET_SR = 16000  # YAMNet expects 16 kHz mono waveforms\n\ndef yamnet_frame_embeddings(path, sr=YAMNET_SR):\n    \"\"\"Returns (n_frames, 1024) -- one embedding per ~0.96s YAMNet frame.\"\"\"\n    y_audio, _ = librosa.load(path, sr=sr, mono=True)\n    waveform = tf.convert_to_tensor(y_audio, dtype=tf.float32)\n    _, embeddings, _ = yamnet_model(waveform)   # scores, embeddings, log-mel spectrogram\n    return embeddings.numpy()\n\nprint(\"Extracting YAMNet frame embeddings for every clip (a few minutes on CPU)...\")\nemb_list, y_emb_list, clip_id_list = [], [], []\nfor clip_idx, (p, label) in enumerate(zip(file_paths, y)):\n    try:\n        frames = yamnet_frame_embeddings(p)\n        if frames.shape[0] == 0:\n            continue\n        emb_list.append(frames)\n        y_emb_list.extend([label] * frames.shape[0])\n        clip_id_list.extend([clip_idx] * frames.shape[0])\n    except Exception as e:\n        print(f\"Skipping for YAMNet (failed to decode): {p} -- {type(e).__name__}: {e}\")\n\nX_yamnet = np.vstack(emb_list)\ny_yamnet = np.array(y_emb_list)\nclip_ids_yamnet = np.array(clip_id_list)\nprint(\n    \"Frame-level embedding matrix:\", X_yamnet.shape,\n    f\"from {len(emb_list)} clips -> {X_yamnet.shape[0] / len(emb_list):.1f} frames/clip on average\"\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:31:41.447460Z","iopub.execute_input":"2026-07-12T09:31:41.447842Z","iopub.status.idle":"2026-07-12T09:35:05.153031Z","shell.execute_reply.started":"2026-07-12T09:31:41.447797Z","shell.execute_reply":"2026-07-12T09:35:05.151922Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Split by CLIP, not by frame -- frames from the same clip must land on the\n# same side of the split, otherwise near-duplicate frames leak between train\n# and test and inflate the score. Same 80/20 stratified methodology as the\n# GRU/CNN split (indices may differ slightly since a handful of clips can\n# fail to decode independently in each pass -- comparable in aggregate,\n# not guaranteed identical row-for-row).\nunique_clip_idx = np.arange(len(file_paths))\nclip_labels = np.array(y)\ntrain_clip_idx, test_clip_idx = train_test_split(\n    unique_clip_idx, test_size=0.2, stratify=clip_labels, random_state=42\n)\ntrain_clip_set = set(train_clip_idx.tolist())\ntest_clip_set = set(test_clip_idx.tolist())\n\ntrain_mask = np.array([cid in train_clip_set for cid in clip_ids_yamnet])\ntest_mask = np.array([cid in test_clip_set for cid in clip_ids_yamnet])\n\nX_yam_train, y_yam_train = X_yamnet[train_mask], y_yamnet[train_mask]\nX_yam_test, y_yam_test = X_yamnet[test_mask], y_yamnet[test_mask]\nclip_ids_test = clip_ids_yamnet[test_mask]\n\nprint(\"Train frames:\", X_yam_train.shape[0], \"| Test frames:\", X_yam_test.shape[0])\n\n# ---- Classifier head on FROZEN embeddings ----\n# A linear head (Logistic Regression) is deliberately simple: YAMNet already\n# did the hard representation-learning work on 2M+ clips, so the head just\n# needs to draw boundaries in that space -- and a simple head is far less\n# likely to overfit the few thousand frame examples we have than another\n# deep network stacked on top.\nscaler_yam = StandardScaler()\nX_yam_train_s = scaler_yam.fit_transform(X_yam_train)\nX_yam_test_s = scaler_yam.transform(X_yam_test)\n\nyamnet_head = LogisticRegression(max_iter=2000, C=1.0, random_state=42)\nyamnet_head.fit(X_yam_train_s, y_yam_train)\n\nframe_pred = yamnet_head.predict(X_yam_test_s)\nframe_acc = accuracy_score(y_yam_test, frame_pred)\nprint(f\"YAMNet + LogReg frame-level accuracy: {frame_acc:.3f} (each frame scored independently)\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:35:05.154172Z","iopub.execute_input":"2026-07-12T09:35:05.154890Z","iopub.status.idle":"2026-07-12T09:37:17.252734Z","shell.execute_reply.started":"2026-07-12T09:35:05.154830Z","shell.execute_reply":"2026-07-12T09:37:17.251892Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Clip-level: average a clip's frame probabilities, then take the argmax --\n# this is what's directly comparable to the GRU/CNN accuracy/macro-F1 above.\nframe_probs = yamnet_head.predict_proba(X_yam_test_s)\nclasses_ = yamnet_head.classes_\n\nclip_true, clip_pred = [], []\nfor cid in sorted(set(clip_ids_test.tolist())):\n    frame_mask = clip_ids_test == cid\n    avg_probs = frame_probs[frame_mask].mean(axis=0)\n    clip_pred.append(classes_[np.argmax(avg_probs)])\n    clip_true.append(y_yam_test[frame_mask][0])\n\nyamnet_acc, yamnet_f1 = report_scores(\n    \"YAMNet embeddings + LogReg (clip-level, soft-vote)\", np.array(clip_true), np.array(clip_pred)\n)\n\nprint(f\"\\n=== From-scratch deep learning vs. transfer learning (clip-level) ===\")\nprint(f\"GRU (from scratch)       | Acc: {gru_acc:.3f} | Macro-F1: {gru_f1:.3f}\")\nprint(f\"CNN (from scratch)       | Acc: {cnn_acc:.3f} | Macro-F1: {cnn_f1:.3f}\")\nprint(f\"YAMNet + LogReg (frozen) | Acc: {yamnet_acc:.3f} | Macro-F1: {yamnet_f1:.3f}\")\nprint(f\"(for reference) LogReg, 37 hand features | Acc: 0.566 | Macro-F1: 0.561\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:37:17.253697Z","iopub.execute_input":"2026-07-12T09:37:17.254025Z","iopub.status.idle":"2026-07-12T09:37:17.330639Z","shell.execute_reply.started":"2026-07-12T09:37:17.253996Z","shell.execute_reply":"2026-07-12T09:37:17.329700Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9. Data Augmentation Alternative: Conditional GAN vs. SMOTE\n\nSMOTE creates synthetic minority samples by interpolating between real neighbors --\nfast and effective, but limited to the convex hull of existing points. A **GAN**\ncan, in principle, learn the actual class-conditional feature distribution and\ngenerate more varied synthetic samples.\n\nHere we use a lightweight **conditional GAN** (cGAN) on the 37-dim handcrafted\nfeature vectors (not raw audio -- a full audio/spectrogram GAN needs far more data\nthan a few hundred clips to avoid mode collapse or noise). The generator takes\nrandom noise **plus a one-hot species label** and learns to produce a plausible\nfeature vector for that species; the discriminator learns to tell real vs.\ngenerated vectors apart for a given label. This is a reasonable fit for a small,\nlow-dimensional, tabular-style feature set like ours -- a much harder ask for\nraw audio waveforms or full spectrogram images at this data scale.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler\n\nLATENT_DIM = 16\nscaler_gan = MinMaxScaler(feature_range=(-1, 1))\nX_scaled = scaler_gan.fit_transform(X_imb)\ny_int_gan = label_enc.transform(y_imb)\n\ndef build_generator(latent_dim, num_classes, n_features):\n    noise_in = layers.Input(shape=(latent_dim,))\n    label_in = layers.Input(shape=(num_classes,))\n    x = layers.Concatenate()([noise_in, label_in])\n    x = layers.Dense(64, activation=\"relu\")(x)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    out = layers.Dense(n_features, activation=\"tanh\")(x)   # tanh matches [-1, 1] scaling\n    return models.Model([noise_in, label_in], out, name=\"generator\")\n\ndef build_discriminator(num_classes, n_features):\n    feat_in = layers.Input(shape=(n_features,))\n    label_in = layers.Input(shape=(num_classes,))\n    x = layers.Concatenate()([feat_in, label_in])\n    x = layers.Dense(64, activation=\"relu\")(x)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    out = layers.Dense(1, activation=\"sigmoid\")(x)\n    return models.Model([feat_in, label_in], out, name=\"discriminator\")\n\nn_features = X_scaled.shape[1]\ngenerator = build_generator(LATENT_DIM, num_classes, n_features)\ndiscriminator = build_discriminator(num_classes, n_features)\ndiscriminator.compile(optimizer=\"adam\", loss=\"binary_crossentropy\", metrics=[\"accuracy\"])\ndiscriminator.trainable = False\n\ngan_noise_in = layers.Input(shape=(LATENT_DIM,))\ngan_label_in = layers.Input(shape=(num_classes,))\ngan_output = discriminator([generator([gan_noise_in, gan_label_in]), gan_label_in])\ngan = models.Model([gan_noise_in, gan_label_in], gan_output, name=\"cgan\")\ngan.compile(optimizer=\"adam\", loss=\"binary_crossentropy\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:37:17.331537Z","iopub.execute_input":"2026-07-12T09:37:17.331868Z","iopub.status.idle":"2026-07-12T09:37:17.412783Z","shell.execute_reply.started":"2026-07-12T09:37:17.331820Z","shell.execute_reply":"2026-07-12T09:37:17.411933Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def one_hot(labels_int, num_classes):\n    return tf.keras.utils.to_categorical(labels_int, num_classes=num_classes)\n\ny_onehot_gan = one_hot(y_int_gan, num_classes)\n\nEPOCHS = 2000\nBATCH_SIZE = 32\nn_samples = X_scaled.shape[0]\n\nfor epoch in range(1, EPOCHS + 1):\n    # --- Train discriminator on a batch of real + fake samples ---\n    idx = np.random.randint(0, n_samples, BATCH_SIZE)\n    real_feats, real_labels = X_scaled[idx], y_onehot_gan[idx]\n\n    noise = np.random.normal(0, 1, (BATCH_SIZE, LATENT_DIM))\n    fake_labels_int = np.random.randint(0, num_classes, BATCH_SIZE)\n    fake_labels = one_hot(fake_labels_int, num_classes)\n    fake_feats = generator.predict([noise, fake_labels], verbose=0)\n\n    discriminator.trainable = True\n    d_loss_real = discriminator.train_on_batch([real_feats, real_labels], np.ones((BATCH_SIZE, 1)))\n    d_loss_fake = discriminator.train_on_batch([fake_feats, fake_labels], np.zeros((BATCH_SIZE, 1)))\n    discriminator.trainable = False\n\n    # --- Train generator via the combined GAN model ---\n    noise = np.random.normal(0, 1, (BATCH_SIZE, LATENT_DIM))\n    gen_labels_int = np.random.randint(0, num_classes, BATCH_SIZE)\n    gen_labels = one_hot(gen_labels_int, num_classes)\n    g_loss = gan.train_on_batch([noise, gen_labels], np.ones((BATCH_SIZE, 1)))\n\n    if epoch % 500 == 0:\n        d_loss = 0.5 * (np.add(d_loss_real, d_loss_fake))\n        print(f\"Epoch {epoch}/{EPOCHS} | D loss: {d_loss[0]:.3f}, D acc: {d_loss[1]:.3f} | G loss: {g_loss:.3f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:37:17.413640Z","iopub.execute_input":"2026-07-12T09:37:17.413926Z","iopub.status.idle":"2026-07-12T09:40:32.831985Z","shell.execute_reply.started":"2026-07-12T09:37:17.413902Z","shell.execute_reply":"2026-07-12T09:40:32.830986Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Use the trained generator to balance classes up to the largest class size,\n# the same target SMOTE would aim for, so the two augmentation strategies\n# are directly comparable.\nclass_counts = Counter(y_int_gan)\nmax_count = max(class_counts.values())\n\nX_gan_parts, y_gan_parts = [X_scaled], [y_int_gan]\nfor cls in range(num_classes):\n    n_needed = max_count - class_counts[cls]\n    if n_needed <= 0:\n        continue\n    noise = np.random.normal(0, 1, (n_needed, LATENT_DIM))\n    labels = one_hot(np.full(n_needed, cls), num_classes)\n    synthetic = generator.predict([noise, labels], verbose=0)\n    X_gan_parts.append(synthetic)\n    y_gan_parts.append(np.full(n_needed, cls))\n\nX_gan_bal = np.vstack(X_gan_parts)\ny_gan_bal_int = np.concatenate(y_gan_parts)\ny_gan_bal = label_enc.inverse_transform(y_gan_bal_int)\n\nprint(\"Balanced dataset (cGAN):\", X_gan_bal.shape, \"Counts:\", Counter(y_gan_bal))\n\n# --- Compare cGAN-balanced data against SMOTE-balanced data with the same CV setup ---\nresults_cv_gan = {}\nprint(\"\\n=== cGAN-balanced, Full (37) feature set ===\")\nfor model_name, model in cv_models.items():\n    scores = cross_validate(model, X_gan_bal, y_gan_bal, cv=skf,\n                             scoring=[\"accuracy\", \"f1_macro\"], n_jobs=-1)\n    acc_mean = scores[\"test_accuracy\"].mean()\n    f1_mean = scores[\"test_f1_macro\"].mean()\n    results_cv_gan[model_name] = (acc_mean, f1_mean)\n    print(f\"{model_name} | Acc: {acc_mean:.3f} | Macro-F1: {f1_mean:.3f}\")\n\nprint(\"\\n=== Side-by-side: SMOTE vs cGAN (Full 37 features, RandomForest/LogReg/KNN/SVM) ===\")\nfor model_name in cv_models:\n    smote_acc, smote_f1 = results_cv_smote[\"Full (37)\"][model_name]\n    gan_acc, gan_f1 = results_cv_gan[model_name]\n    print(f\"{model_name:14s} | SMOTE Acc/F1: {smote_acc:.3f}/{smote_f1:.3f} \"\n          f\"| cGAN Acc/F1: {gan_acc:.3f}/{gan_f1:.3f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:40:32.833005Z","iopub.execute_input":"2026-07-12T09:40:32.833330Z","iopub.status.idle":"2026-07-12T09:40:37.058224Z","shell.execute_reply.started":"2026-07-12T09:40:32.833301Z","shell.execute_reply":"2026-07-12T09:40:37.057384Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Reading the GAN result:** with only a few hundred real samples, a GAN this small\nis prone to producing less diverse (or noisier) synthetic points than SMOTE's\nneighbor-interpolation approach -- don't be surprised if SMOTE matches or beats it\nhere. GANs tend to pull ahead once there's enough real data per class (low\nhundreds to thousands) for the generator to learn a meaningful distribution\ninstead of memorizing/collapsing. Worth revisiting once `NUM_SPECIES` and sample\ncounts are scaled up.\n\nThis comparison now runs on the synthetically imbalanced `X_imb`/`y_imb` built at the start of Section 7 (two species downsampled to `MINORITY_SIZE`), not the naturally-balanced full dataset -- otherwise SMOTE and the cGAN would both have nothing to correct and trivially return identical results, which is what happened before this fix.","metadata":{}},{"cell_type":"markdown","source":"## 9b. Does the SMOTE vs. cGAN Gap Close With More Minority Data?\n\nThe Section 9 result (SMOTE beating the cGAN at `MINORITY_SIZE = 15`) matches\nthe caveat stated above -- but that caveat was, until now, just an assertion.\nHere we actually test it: rerun the same SMOTE-vs-cGAN comparison at a few\ndifferent minority-class sizes (`15 -> 40 -> 70` real samples) and track\nwhether the accuracy gap between the two shrinks as the cGAN's generator gets\nmore real data per class to learn from.\n\nTo keep the sweep's runtime reasonable, the cGAN uses **1,000 training epochs\nper minority size** here instead of Section 9's 2,000 (same architecture,\nsame training logic -- just half the epochs, run three times instead of\nonce). Absolute numbers may shift slightly from Section 9's `MINORITY_SIZE =\n15` case as a result; the trend across sizes is what matters.","metadata":{}},{"cell_type":"code","source":"def build_imbalanced_data(X_full, y_full, minority_cls, minority_size, seed=42):\n    rng = np.random.RandomState(seed)\n    idx_parts = []\n    for cls in np.unique(y_full):\n        cls_idx = np.where(y_full == cls)[0]\n        if cls in minority_cls:\n            cls_idx = rng.choice(cls_idx, size=min(minority_size, len(cls_idx)), replace=False)\n        idx_parts.append(cls_idx)\n    idx = np.concatenate(idx_parts)\n    rng.shuffle(idx)\n    return X_full[idx], y_full[idx]\n\n\ndef train_cgan_quick(X_scaled_in, y_int_in, n_classes, latent_dim=LATENT_DIM,\n                      epochs=1000, batch_size=32, seed=42):\n    \"\"\"Lighter version of the Section 9 training loop (fewer epochs, run\n    silently) so the 3-point sweep finishes in reasonable time. Same\n    generator/discriminator architecture (build_generator/build_discriminator\n    from Section 9), same alternating-training logic.\"\"\"\n    tf.random.set_seed(seed)\n    gen = build_generator(latent_dim, n_classes, X_scaled_in.shape[1])\n    disc = build_discriminator(n_classes, X_scaled_in.shape[1])\n    disc.compile(optimizer=\"adam\", loss=\"binary_crossentropy\", metrics=[\"accuracy\"])\n    disc.trainable = False\n\n    noise_in = layers.Input(shape=(latent_dim,))\n    label_in = layers.Input(shape=(n_classes,))\n    gan_out = disc([gen([noise_in, label_in]), label_in])\n    cgan = models.Model([noise_in, label_in], gan_out)\n    cgan.compile(optimizer=\"adam\", loss=\"binary_crossentropy\")\n\n    y_oh = one_hot(y_int_in, n_classes)\n    n = X_scaled_in.shape[0]\n    for _ in range(1, epochs + 1):\n        idx = np.random.randint(0, n, batch_size)\n        real_feats, real_labels = X_scaled_in[idx], y_oh[idx]\n        noise = np.random.normal(0, 1, (batch_size, latent_dim))\n        fake_labels_int = np.random.randint(0, n_classes, batch_size)\n        fake_labels = one_hot(fake_labels_int, n_classes)\n        fake_feats = gen.predict([noise, fake_labels], verbose=0)\n\n        disc.trainable = True\n        disc.train_on_batch([real_feats, real_labels], np.ones((batch_size, 1)))\n        disc.train_on_batch([fake_feats, fake_labels], np.zeros((batch_size, 1)))\n        disc.trainable = False\n\n        noise = np.random.normal(0, 1, (batch_size, latent_dim))\n        gen_labels_int = np.random.randint(0, n_classes, batch_size)\n        gen_labels = one_hot(gen_labels_int, n_classes)\n        cgan.train_on_batch([noise, gen_labels], np.ones((batch_size, 1)))\n    return gen\n\n\ndef cgan_balance(gen, X_scaled_in, y_int_in, n_classes, latent_dim=LATENT_DIM):\n    class_counts = Counter(y_int_in)\n    max_count = max(class_counts.values())\n    X_parts, y_parts = [X_scaled_in], [y_int_in]\n    for cls in range(n_classes):\n        n_needed = max_count - class_counts[cls]\n        if n_needed <= 0:\n            continue\n        noise = np.random.normal(0, 1, (n_needed, latent_dim))\n        labels = one_hot(np.full(n_needed, cls), n_classes)\n        synthetic = gen.predict([noise, labels], verbose=0)\n        X_parts.append(synthetic)\n        y_parts.append(np.full(n_needed, cls))\n    return np.vstack(X_parts), np.concatenate(y_parts)\n\n\ndef eval_balanced(X_bal_in, y_bal_in, models_dict, cv):\n    out = {}\n    for name, model in models_dict.items():\n        scores = cross_validate(model, X_bal_in, y_bal_in, cv=cv,\n                                 scoring=[\"accuracy\", \"f1_macro\"], n_jobs=-1)\n        out[name] = (scores[\"test_accuracy\"].mean(), scores[\"test_f1_macro\"].mean())\n    return out\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:40:37.059147Z","iopub.execute_input":"2026-07-12T09:40:37.059463Z","iopub.status.idle":"2026-07-12T09:40:37.076271Z","shell.execute_reply.started":"2026-07-12T09:40:37.059434Z","shell.execute_reply":"2026-07-12T09:40:37.075376Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"MINORITY_SIZES = [15, 40, 70]\nsweep_skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\nsweep_avg_gap = []   # mean(SMOTE acc - cGAN acc) across the 4 models, per minority size\n\nfor m_size in MINORITY_SIZES:\n    print(f\"\\n########## MINORITY_SIZE = {m_size} ##########\")\n    X_sw, y_sw = build_imbalanced_data(X, y, minority_species, m_size)\n    print(\"Imbalanced counts:\", Counter(y_sw))\n\n    # --- SMOTE side ---\n    min_cls_size = min(Counter(y_sw).values())\n    smote_sw = SMOTE(random_state=42, k_neighbors=min(5, max(1, min_cls_size - 1)))\n    X_bal_sw, y_bal_sw = smote_sw.fit_resample(X_sw, y_sw)\n    smote_scores = eval_balanced(X_bal_sw, y_bal_sw, cv_models, sweep_skf)\n\n    # --- cGAN side ---\n    scaler_sw = MinMaxScaler(feature_range=(-1, 1))\n    X_sw_scaled = scaler_sw.fit_transform(X_sw)\n    y_int_sw = label_enc.transform(y_sw)\n    gen_sw = train_cgan_quick(X_sw_scaled, y_int_sw, num_classes, epochs=1000)\n    X_gan_sw, y_gan_int_sw = cgan_balance(gen_sw, X_sw_scaled, y_int_sw, num_classes)\n    y_gan_sw = label_enc.inverse_transform(y_gan_int_sw)\n    cgan_scores = eval_balanced(X_gan_sw, y_gan_sw, cv_models, sweep_skf)\n\n    print(f\"{'Model':14s} | {'SMOTE Acc':>10s} | {'cGAN Acc':>10s} | {'Gap (SMOTE-cGAN)':>18s}\")\n    gaps = []\n    for name in cv_models:\n        s_acc, g_acc = smote_scores[name][0], cgan_scores[name][0]\n        gap = s_acc - g_acc\n        gaps.append(gap)\n        print(f\"{name:14s} | {s_acc:10.3f} | {g_acc:10.3f} | {gap:18.3f}\")\n\n    sweep_avg_gap.append(np.mean(gaps))\n\nprint(\"\\n=== Average SMOTE-minus-cGAN accuracy gap by MINORITY_SIZE ===\")\nfor m_size, gap in zip(MINORITY_SIZES, sweep_avg_gap):\n    verdict = \"SMOTE ahead\" if gap > 0.01 else \"cGAN ahead\" if gap < -0.01 else \"roughly tied\"\n    print(f\"MINORITY_SIZE={m_size:3d} | avg gap: {gap:+.3f}  ({verdict})\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:40:37.077238Z","iopub.execute_input":"2026-07-12T09:40:37.077533Z","iopub.status.idle":"2026-07-12T09:45:56.843486Z","shell.execute_reply.started":"2026-07-12T09:40:37.077503Z","shell.execute_reply":"2026-07-12T09:45:56.842541Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(7, 4))\nplt.axhline(0, color=\"gray\", linewidth=0.8, linestyle=\"--\")\nplt.plot(MINORITY_SIZES, sweep_avg_gap, marker=\"o\", color=\"darkorange\")\nplt.title(\"SMOTE vs. cGAN: accuracy gap by minority-class size\")\nplt.xlabel(\"Minority class size (real samples per class)\")\nplt.ylabel(\"Avg accuracy gap\\n(SMOTE - cGAN; positive = SMOTE ahead)\")\nplt.xticks(MINORITY_SIZES)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-12T09:45:56.844530Z","iopub.execute_input":"2026-07-12T09:45:56.845062Z","iopub.status.idle":"2026-07-12T09:45:57.020981Z","shell.execute_reply.started":"2026-07-12T09:45:56.845033Z","shell.execute_reply":"2026-07-12T09:45:57.020161Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Reading the sweep:** if the gap shrinks (moves toward zero) or flips\nnegative as `MINORITY_SIZE` grows, that confirms the notebook's earlier\nclaim -- the cGAN needed more real examples per class to learn a useful\nclass-conditional distribution, and SMOTE's advantage was a small-data\nartifact, not a fundamental one. If the gap stays flat or widens instead,\nthat's evidence this cGAN architecture (or its fixed 1,000-1,000-2,000-epoch\nbudget) is the bottleneck rather than data volume alone -- worth revisiting\nwith a larger/deeper generator or longer training before drawing conclusions\neither way.","metadata":{}},{"cell_type":"markdown","source":"## 10. Summary & Next Steps\n\n**Implemented in this notebook:**\n- A 37-dimensional handcrafted acoustic feature vector (MFCC + spectral + chroma + ZCR) per recording.\n- Feature selection via ANOVA F-test and Random Forest importance, with an explicit overlap analysis.\n- Classical ML baselines (Random Forest, Logistic Regression, KNN, SVM) on full vs. reduced feature sets, reporting **accuracy and macro-F1**.\n- Class balancing via **SMOTE**, validated with stratified 5-fold cross-validation.\n- A **GRU** sequence model trained directly on mel-spectrogram time steps, as an alternative to handcrafted features.\n- A true **2D CNN** on mel-spectrogram images (Conv2D + pooling blocks) with mild **SpecAugment**-style time/frequency masking on the training split, directly comparable to the GRU on the same train/test split.\n- **Transfer learning with YAMNet** (AudioSet-pretrained) embeddings, frame-level (not clip-averaged) to multiply the effective training set size, with a frozen-embedding Logistic Regression head and clip-level soft-voting for evaluation -- this is the approach that meaningfully outperforms the from-scratch GRU/CNN on a dataset this small.\n- A **conditional GAN** trained to generate synthetic minority-class feature vectors, benchmarked head-to-head against SMOTE -- on a synthetically imbalanced copy of the data, since the dataset is balanced by construction and neither method has anything to correct otherwise.\n- A **MINORITY_SIZE sweep (15 -> 40 -> 70)** re-running the SMOTE-vs-cGAN comparison at increasing minority-class sample counts, to test (rather than just assert) whether the cGAN's disadvantage shrinks with more real data per class.\n- Full **`GridSearchCV`/`RandomizedSearchCV`** hyperparameter tuning for all four classical models (RandomForest, SVM tuned via `RandomizedSearchCV`; LogReg, KNN via `GridSearchCV`), evaluated against the untuned baselines on the same held-out split.\n- A configurable `NUM_SPECIES` knob so the whole pipeline scales to more species once enough recordings per class are available.\n\n**Still open:**\n- Re-running the whole pipeline (baselines, deep models, GAN comparison) at a larger `NUM_SPECIES`, where every classical/deep model has more classes and more total data to work with.\n- Tuning the CNN/GRU/cGAN's own hyperparameters (architecture size, learning rate, epoch budget) the same systematic way Section 6b now tunes the classical models -- currently still hand-picked.\n- Ensembling the tuned classical models with the YAMNet transfer-learning head, since they make different kinds of errors and a simple vote/average may beat any single model.\n","metadata":{}}]}