{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","file_extension":".py"},"accelerator":"GPU"},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"68db111f","cell_type":"markdown","source":"# ARGUS — Multi-Species Few-Shot Bioacoustic Detection\n\nKaggle notebook, built from the tested pipeline (`arguswala_setup_v2.py`,\n`arguswala_v2.py`, `arguswala_sweeps.py`, `arguswala_ladder.py`, `arguswala_confounds.py`,\n`arguswala_distill.py`, `arguswala_frontier.py` — 371 local assertions passing as of the\nlast edit; see `_test_v2_logic.py` / `_test_ladder.py` / `_test_confounds.py` /\n`_test_distill.py` / `_test_frontier.py` in the project repo). One notebook cell per\nsource file, run top to bottom.\n\n## Current state of this build (31 Aug 2026) — FRONTIER + WILD-SCAN SESSION\n\n`SMOKE_TEST = False`. **Do not set it True** — it cuts `TARGET_LIST` to 2 species and\nbreaks the `[:5]` slices in Cells 5–7.\n\n**This build runs Cell 7 only, and spends a full session on two open questions.**\nProjected **~5.5 h** against the 12 h ceiling, with both parts checkpointed.\n\n**Part A — bottleneck frontier (~3.5 h).** Cell 6 distilled Perch into the stage-1 network\nand lost 0.136 AUC. This tests the most specific explanation for *why*: Perch emits 1536\ndimensions, the student's trunk emits 128, and a linear head from 128-d can only reproduce\na 128-dimensional subspace of Perch's space no matter how long it trains. Sweeping the\nbottleneck (64 / 128 / 256 / 512, plus a nonlinear head at 512) turns a single pass/fail\ninto an accuracy-vs-size curve — 25× down to 7× smaller than Perch. All three outcomes are\npre-registered in the cell header, including the awkward one where fidelity rises but\nspecies AUC does not.\n\n**Part B — wild-audio scan (~2 h).** The pipeline has never been asked to find a call a\nreal bird actually made; every number so far comes from calls *planted* into soundscape.\nThis scans 500 unlabelled soundscapes (~33 h of audio), strictly disjoint from the 14\ncalibration/evaluation beds, at the same best-F1 threshold everything else uses — plus the\nsame scan for a control species on identical files. Saves the top 60 detections per species\nas listenable clips. It cannot *confirm* a detection (no ground truth, no expert ear), but\nit establishes a rate, a ranked shortlist, and a comparison.\n\nFlag state: `SKIP_MAIN_LOOP` / `SKIP_LADDER` / `SKIP_SWEEPS` / `SKIP_DISTILL` **`True`**;\nCell 5 `RUN_P1`/`RUN_P2`/`RUN_P3` **`False`**; Cell 7 `SKIP_FRONTIER` **`False`**.\nCell 6 skips but still defines `build_distill_set` and `make_distill_embedder`, which\nCell 7 reuses — so it must still execute.\n\n**Session log:**\n\n| Session | Status | Result |\n|---|---|---|\n| P2 re-run, 6 seeds | done 29 Aug, 3h40m | Settled: verified **falls −0.050**, negative 5/5. |\n| Widened sweeps, 5sp×3seed | done 30 Aug, 8h13m | Two conclusions strengthened, shot-count gain shrank 0.116→0.078. |\n| Perch distillation | done 30 Aug, 1h26m | −0.136 mean; free for one species, chance-level for the study species. |\n| Frontier + wild scan | **this build**, ~5.5 h | — |\n| DCASE external validation | not started | Roadmap §4, 5-day timebox. |\n\n## Before running — checklist\n\n1. **Notebook settings (right sidebar):** Accelerator = **GPU**, Internet = **On**.\n2. **Add data**, attach:\n   - **BirdCLEF 2024** competition data (`birdclef-2024`)\n   - **A Perch model.** Which one depends on `PERCH_VERSION` in Cell 1 (default `\"v2\"`):\n     - `PERCH_VERSION = \"v2\"` → attach `google/bird-vocalization-classifier` →\n       variation **`perch_v2`** (GPU build) — the CPU fallback `perch_v2_cpu` also works\n       if no GPU is attached, but detection is automatic either way.\n     - `PERCH_VERSION = \"v1\"` → attach `google/bird-vocalization-classifier` →\n       variation **`bird-vocalization-classifier`**, version **8**.\n   - A version mismatch between the attached model and `PERCH_VERSION` raises\n     immediately with a clear message rather than silently scoring the wrong model.\n3. Cells 1–7 build on each other's names — run in order, don't skip cells manually (the\n   skip *flags* inside Cells 2–6 handle that for you; the cells themselves still need to\n   execute so their functions and cached state exist for Cell 7).\n4. Watch Cell 1's first print line — `train_metadata.csv columns: [...]` — and confirm\n   `author` is listed (Cell 5's recordist-disjoint test needs it; already confirmed\n   present in the 27 Aug run, but re-check if the attached dataset version ever changes).\n5. Watch Cell 5's own timing print after the very first (species, seed) pair completes —\n   it projects a total for that section. Stop and widen `PERCH_ALONE_STRIDE_S` if it\n   looks too slow rather than finding out hours later.\n\nSee `ARGUS_Roadmap_Aug_Sept_2026.md` (§1b) for why the ablation ladder (Cell 4) and a\nPerch-version decision should both happen on a *few* species before committing to the\nfull multi-species run in Cell 2 — decoys are ranked by Perch-embedding proximity, so\nchanging the verifier config after Cell 2 has started makes species runs incomparable.\n","metadata":{}},{"id":"fd1e21ff","cell_type":"markdown","source":"## Cell 1 — Setup\n\n`arguswala_setup_v2.py`","metadata":{}},{"id":"2dae41c1","cell_type":"code","source":"# ===== CELL 1 v2 / SETUP -- multi-species, shot-strategy aware =====\n# Replaces arguswala_setup.py. Run BEFORE arguswala_v2.py.\n#\n# WHAT CHANGED vs v1, and why:\n#\n#  (1) MULTI-TARGET. v1 hard-wired one TARGET. Everything per-species is now behind\n#      build_target(sp), so the same pipeline evaluates N species. This pays the debt\n#      arguswala.py:375 admits to (\"one target species is not a validated finding\").\n#\n#  (2) SPECIES CACHE COMPUTED ONCE. v1 embedded 30 candidate species per target. With\n#      10 targets that is 10x redundant Perch inference -- the dominant cost. The cache\n#      is now global; build_target() only ranks against it. ~10x saving on setup.\n#\n#  (3) SHOT-SELECTION STRATEGY is now a variable, not an assumption. v1 always took the\n#      5 highest-`rating` recordings. DCASE always takes the first 5. Nolasco et al.\n#      (2023) sec.4.1 report that unrepresentative support sets were the dominant failure\n#      cause their expert annotators identified, and call better selection an open\n#      problem. So it becomes an axis we measure instead of a default we inherit.\n#\n#  (4) PER-SPECIES COVARIATES are recorded at build time (decoy proximity, stereotypy,\n#      spectral centre/bandwidth, call duration, n recordings). Nolasco et al. tried\n#      multivariate regression to find what predicts task difficulty and FAILED to find\n#      any consistent factor. These covariates are free at runtime and impossible to\n#      reconstruct afterwards -- logging them is what makes that analysis possible later.\n#\n#  (5) The held-out pool is now derived from the shots actually chosen, not from a\n#      positional slice. Under a non-'first' strategy, files[N_SHOT:] would have leaked\n#      support recordings into the evaluation pool.\nimport os, glob, time\nimport numpy as np\nimport pandas as pd\nimport librosa\nimport tensorflow_hub as hub\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.cluster import KMeans\n\n# ---------------------------------------------------------------- ablation ladder\n# Each rung is independently switchable so marginal contributions are measurable.\n# Rung 0 is the original pipeline; nothing below changes it when all flags are False.\nVERIFIER = \"logreg\"        # 'logreg' (rung 0) | 'proto' (rung 1, stage 2)\n\n# ------------------------------------------------------------------ configuration\nSR = 22050\nSEG_SEC = 1.0\nCLEN = int(SEG_SEC * SR)\n\nN_SHOT          = 5\nSHOT_STRATEGY   = \"rated\"     # 'rated' | 'first' | 'random' | 'diverse'\nN_DECOYS        = 4\nN_CANDIDATES    = 30          # non-target species considered for decoy ranking\nCALIB_PER_SPECIES = 6\nDECOY_RECORDINGS  = 20\nN_CALIB_BEDS, N_EVAL_BEDS = 4, 10\nSNR_LEVELS = (9, 6, 3, 0, -3, -6)\n\n# Targets are PRE-REGISTERED: chosen by a rule, before any result is seen.\n# Rule: species with >= MIN_RECS recordings, ranked by recording count, top N_TARGETS,\n# with the original study species forced in first. Deterministic -- no cherry-picking.\nPRIMARY_TARGET = \"whbsho3\"    # White-bellied Sholakili, the original study species\nN_TARGETS      = 10\nMIN_RECS       = N_SHOT + 4   # 5 shots + >=4 genuinely held out\n\nSMOKE_TEST = False            # 15 Aug 2026: timing measured (10 min/species/3seeds),\n                              # config locked (PCEN/logreg/no-tfa/no-hn all confirmed\n                              # winners in the ladder), noise floor known (0.025-0.035).\n                              # Centrepiece run: full N_TARGETS species, no more reason\n                              # to stay on the 2-species smoke test.\n\n# Added 14 Aug 2026, after the first real run: whbsho3 stage-1-alone AUC came back 0.517\n# in a run that excludes 2 species from the encoder bank (this multi-target pipeline),\n# vs. the original single-species baseline of 0.459, which only ever excluded whbsho3.\n# ENCODER_SEEDS (widened separately, in CELL 2) tests training-run-to-training-run noise\n# on a FIXED bank. This flag tests the OTHER candidate cause: does the bank's COMPOSITION\n# itself change when only 1 species is excluded instead of 2? Set True to force\n# TARGET_LIST down to [PRIMARY_TARGET] alone -- the exact exclusion set the original\n# baseline used -- so build_banks() in CELL 2 produces a directly comparable bank.\n# Re-run CELL 1 then CELL 2 after flipping this; no kernel restart needed, everything\n# downstream is rebuilt from these globals. CELL 3/4 are not needed for this comparison.\nSINGLE_TARGET_COMPARISON = False\n\n# Added 29 Aug 2026 for CELL 6 (Perch distillation). Selects which embedding space the\n# STAGE-2 VERIFIER is trained and run in:\n#   \"perch\"     -- Perch v2's 1536-d embeddings. Every result published before this flag\n#                  existed used this, and it stays the default.\n#   \"distilled\" -- a ~445k-param CNN trained (CELL 6) to reproduce Perch's embeddings.\n#                  At detection time NO Perch inference runs at all.\n#\n# DELIBERATELY NARROW. Only the verifier's embeddings switch. Decoy ranking\n# (SPECIES_CACHE, perch_embed) and 'diverse' shot selection stay on Perch under BOTH\n# settings, because those choose WHAT THE TEST IS -- which species count as decoys, and\n# which five calls are the support set. If the student picked its own decoys, the two arms\n# would be scored against different test sets and the comparison would be meaningless.\n# This is the same trap the literature review flagged for the Perch v1->v2 upgrade\n# (\"change Perch, and the decoy set changes, and the difficulty of the specificity test\n# changes with it\"). Holding the test fixed is what makes teacher-vs-student a clean A/B.\nEMBED_BACKEND = \"perch\"\nDISTILL_EMBED_CTX = None      # set by CELL 6 to the trained student's embedder\n\ndef verifier_embed_ctx(y22, centres_sec):\n    \"\"\"Embeddings the stage-2 verifier is FIT on and SCORED with. y22 is a 22050 Hz\n    recording; centres_sec are event centres in seconds. Dispatches on EMBED_BACKEND so\n    that build_target()'s verifier fit and CELL 2's detect_pv() always agree about which\n    embedding space they are in -- a mismatch there would fit the classifier in one space\n    and score candidates in another, which produces plausible-looking nonsense rather\n    than an error.\"\"\"\n    if EMBED_BACKEND == \"perch\":\n        return perch_embed_ctx(to32k(y22), centres_sec)\n    if EMBED_BACKEND == \"distilled\":\n        assert DISTILL_EMBED_CTX is not None, \\\n            \"EMBED_BACKEND='distilled' but CELL 6 has not set DISTILL_EMBED_CTX\"\n        return DISTILL_EMBED_CTX(y22, centres_sec)\n    raise ValueError(f\"unknown EMBED_BACKEND {EMBED_BACKEND!r}\")\n\nroot = os.path.dirname(glob.glob('/kaggle/input/**/train_metadata.csv', recursive=True)[0])\nAUD = os.path.join(root, 'train_audio')\nSS  = os.path.join(root, 'unlabeled_soundscapes')\nmeta = pd.read_csv(os.path.join(root, 'train_metadata.csv'))\n\n# Recordist identity column, for the recordist-disjoint confound check (CELL 5, roadmap\n# \"owed\" list: rules out \"did it learn the microphone, not the bird\"). BirdCLEF inherits\n# Xeno-Canto's schema, where the recordist's name is the 'author' field -- but that is a\n# CONVENTION, not something this script has verified against the actual attached dataset,\n# so it is printed here rather than assumed silently. If build_target(recordist_disjoint=\n# True) raises a KeyError, check this list against the real column name and fix RECORDIST_COL.\nRECORDIST_COL = \"author\"\nprint(f\"train_metadata.csv columns: {meta.columns.tolist()}\")\nprint(f\"  RECORDIST_COL = {RECORDIST_COL!r} \"\n      f\"({'present' if RECORDIST_COL in meta.columns else 'NOT FOUND -- fix before using recordist_disjoint'})\")\n\n# ------------------------------------------------------------------ audio helpers\ndef rms(x):\n    return float(np.sqrt(np.mean(np.square(x, dtype=np.float64)) + 1e-12))\n\ndef ld(p, dur=None):\n    y, _ = librosa.load(p, sr=SR, mono=True, duration=dur)\n    return y.astype('float32')\n\ndef l2n(x):\n    x = np.asarray(x, dtype=\"float32\")\n    return x / (np.linalg.norm(x, axis=-1, keepdims=True) + 1e-9)\n\ndef cos(a, b):\n    return float(a @ b / ((np.linalg.norm(a) + 1e-9) * (np.linalg.norm(b) + 1e-9)))\n\ndef loudest_offset(y, n=CLEN):\n    if len(y) <= n: return 0\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    return int(np.argmax(cs[n:] - cs[:-n]))\n\ndef peak_windows(y, k=1, n=CLEN):\n    \"\"\"Top-k non-overlapping highest-energy windows of length n.\"\"\"\n    if len(y) <= n: return [np.pad(y, (0, n - len(y))).astype('float32')]\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    w = cs[n:] - cs[:-n]\n    out, taken = [], []\n    for _ in range(k):\n        pick = next((int(i) for i in np.argsort(w)[::-1] if all(abs(int(i) - t) >= n for t in taken)), None)\n        if pick is None: break\n        taken.append(pick); out.append(y[pick:pick + n].astype('float32'))\n    return out\n\ndef build_support_views(recordings, shifts=(-0.25, 0.0, 0.25)):\n    out = []\n    for y in recordings:\n        base = loudest_offset(y)\n        for shift in shifts:\n            a = int(np.clip(base + shift * SR, 0, max(0, len(y) - CLEN)))\n            w = y[a:a + CLEN]\n            if len(w) < CLEN: w = np.pad(w, (0, CLEN - len(w)))\n            out.append(w.astype('float32'))\n    return out\n\ndef place(y, pool, rng, spans, snr_db, gap_s=0.4):\n    \"\"\"Mix one random call from `pool` into `y` at a free position -> centre in seconds.\"\"\"\n    for _ in range(80):\n        pos = int(rng.integers(0, len(y) - CLEN))\n        if any(pos < e + int(gap_s*SR) and pos + CLEN > s - int(gap_s*SR) for s, e in spans):\n            continue\n        c = pool[int(rng.integers(len(pool)))][:CLEN]\n        if len(c) < CLEN: c = np.pad(c, (0, CLEN - len(c)))\n        g = rms(y[pos:pos+CLEN]) * (10 ** (float(snr_db)/20)) / (rms(c) + 1e-9)\n        y[pos:pos+CLEN] += c.astype('float32') * g\n        spans.append((pos, pos + CLEN))\n        return (pos + CLEN/2) / SR\n    return None\n\n# ------------------------------------------------------------------ Perch\n# PERCH 2.0 UPGRADE (Aug 2025 release, arXiv 2508.04665). Three things changed and each\n# one silently breaks the v1 code path, so all three are handled explicitly. Verified\n# against the official loader: google-research/perch-hoplite, zoo/taxonomy_model_tf.py.\n#\n#  (a) NO `infer_tf`. Perch 1.0 was a TF-Jax export exposing model.infer_tf(). Perch 2.0\n#      exposes only the standard serving signature. perch-hoplite branches on exactly\n#      this: `elif hasattr(self.model, 'infer_tf') ... else: # Perch v2 style\n#      outputs = self.model.signatures['serving_default'](inputs=rebatched_audio)`.\n#      Calling infer_tf on v2 raises AttributeError.\n#\n#  (b) EMBEDDING WIDTH 1280 -> 1536 (EfficientNet-B1 -> B3). Probed, never assumed.\n#      v2 also emits a `spatial_embedding` (5,3,1536) head that must be ignored --\n#      `embedding` is the pooled mean vector we want.\n#\n#  (c) AUDIO IS PEAK-NORMALISED BEFORE INFERENCE. The official wrapper removes DC offset\n#      and scales each window to a peak of 0.25 (zoo_interface.normalize_audio,\n#      target_peak=0.25). ARGUS v1 fed raw waveforms to Perch and did NEITHER -- so\n#      Perch was being given amplitudes outside the distribution it is served with.\n#      This affects the EXISTING v1 numbers too, not just the upgrade.\n#      Note it does NOT disturb the SNR sweep: peak norm is a per-window scalar multiply,\n#      and SNR is a ratio within the window, so planted-call SNR is preserved exactly.\nPERCH_VERSION   = \"v2\"     # \"v2\" | \"v1\"  -- recorded in the results CSV\nPERCH_PEAK_NORM = 0.25     # None to disable (v1 behaviour); 0.25 matches the official wrapper\nPERCH_BACKEND   = \"tf\"     # \"tf\" | \"onnx\" -- see below. Default \"tf\" -> ZERO behaviour\n                           # change to the tested GPU pipeline unless explicitly opted in.\n\nPERCH_SR, PERCH_WIN = 32000, 5.0     # unchanged between versions\nPERCH_LEN = int(PERCH_WIN * PERCH_SR)\nPERCH_BATCH = 16\n\n_V2_SLUGS = (\"perch_v2\", \"perch_v2_cpu\")\ndef _find_local_perch(prefer_v2):\n    \"\"\"Kaggle attaches models as input dirs. Pick one matching the requested version.\"\"\"\n    hits = [os.path.dirname(p) for p in glob.glob('/kaggle/input/**/saved_model.pb', recursive=True)\n            if 'bird-vocalization' in p.lower() or 'perch' in p.lower()]\n    if not hits: return None\n    is_v2 = lambda p: any(s in p.lower() for s in _V2_SLUGS)\n    pool = [p for p in hits if is_v2(p) == prefer_v2] or hits\n    def _ver(p):\n        try: return int(os.path.basename(p))\n        except ValueError: return -1\n    return max(pool, key=_ver)\n\ndef _remote_handle():\n    if PERCH_VERSION != \"v2\":\n        # perch_8 in perch-hoplite's config; v8 is the last 1280-d release.\n        return (\"https://www.kaggle.com/models/google/bird-vocalization-classifier\"\n                \"/tensorFlow2/bird-vocalization-classifier/8\")\n    # perch-hoplite ships separate GPU and CPU builds at different version numbers.\n    try:\n        import tensorflow as tf\n        gpu = bool(tf.config.list_physical_devices('GPU'))\n    except Exception:\n        gpu = False\n    slug, ver = (\"perch_v2\", 2) if gpu else (\"perch_v2_cpu\", 1)\n    return (\"https://www.kaggle.com/models/google/bird-vocalization-classifier\"\n            f\"/tensorFlow2/{slug}/{ver}\")\n\n# ---------------------------------------------------------------- ONNX backend (opt-in)\n# For CPU-constrained / edge-hardware experiments (roadmap: ARGUS_Hardware_Deployment_\n# Roadmap.md), not for the main GPU pipeline. Uses a pre-converted, community-published,\n# Apache-2.0 ONNX export (justinchuby/Perch-onnx on Hugging Face) rather than converting\n# ourselves -- same I/O signature, independently verified against the real 413MB weights\n# on 13 Aug 2026: input 'inputs' [batch,160000], output 'embedding' [batch,1536], real\n# inference measured at ~156ms/clip (batch=16, 4 CPU threads) on a laptop CPU -- NOT the\n# target embedded device, a reference point only. Peak-norm was confirmed to materially\n# change the real embedding (mean abs delta 0.065), i.e. the v1 peak-norm bug was real,\n# not theoretical, on the actual weights.\n_ONNX_HF_URL = \"https://huggingface.co/justinchuby/Perch-onnx/resolve/main/perch_v2_no_dft.onnx\"\n\ndef _find_local_onnx():\n    hits = (glob.glob('/kaggle/input/**/perch_v2_no_dft*.onnx', recursive=True)\n            or glob.glob('/kaggle/input/**/perch_v2*.onnx', recursive=True))\n    return hits[0] if hits else None\n\ndef _load_onnx_perch():\n    import onnxruntime as ort\n    path = _find_local_onnx()\n    if path is None:\n        # Falls back to a live download (~413MB) if nothing is attached as a Kaggle\n        # input -- slow and wasteful on repeated runs. Attaching a Kaggle dataset\n        # mirror (e.g. tuckerarrants/perch-v2-no-dft-onnx) is the fast path.\n        import requests\n        path = \"/kaggle/working/perch_v2_no_dft.onnx\"\n        print(f\"  no local ONNX Perch found; downloading from {_ONNX_HF_URL} ...\")\n        r = requests.get(_ONNX_HF_URL, stream=True, timeout=300)\n        with open(path, \"wb\") as f:\n            for chunk in r.iter_content(chunk_size=1 << 20):\n                f.write(chunk)\n    so = ort.SessionOptions(); so.intra_op_num_threads = 4\n    sess = ort.InferenceSession(path, sess_options=so, providers=[\"CPUExecutionProvider\"])\n    in_name = sess.get_inputs()[0].name\n    out_names = [o.name for o in sess.get_outputs()]\n    print(f\"  ONNX Perch loaded from {path}  |  input='{in_name}'  outputs={out_names}\")\n    return sess, in_name, out_names\n\nif PERCH_BACKEND == \"onnx\":\n    _ort_sess, _ort_in, _ort_out_names = _load_onnx_perch()\n    _HAS_INFER_TF = False; _SERVING = None; perch = None   # not used on this path\n    print(f\"  calling convention: onnxruntime (CPU)  |  peak-norm: {PERCH_PEAK_NORM}\")\nelse:\n    PERCH_HANDLE = _find_local_perch(PERCH_VERSION == \"v2\") or _remote_handle()\n    print(f\"Perch {PERCH_VERSION} source: {PERCH_HANDLE}\")\n    perch = hub.load(PERCH_HANDLE)\n\n    # Which calling convention does THIS artefact expose? Detected, not assumed -- so a\n    # mis-attached Kaggle input degrades to a clear message instead of a confusing crash.\n    _HAS_INFER_TF = hasattr(perch, \"infer_tf\")\n    _SERVING = None\n    if not _HAS_INFER_TF:\n        try:\n            _SERVING = perch.signatures[\"serving_default\"]\n        except Exception as ex:\n            raise RuntimeError(\n                \"Loaded Perch artefact exposes neither infer_tf nor a 'serving_default' \"\n                f\"signature ({type(ex).__name__}). Check the attached Kaggle model.\"\n            ) from ex\n    print(f\"  calling convention: {'infer_tf (Perch 1.x)' if _HAS_INFER_TF else 'serving_default (Perch 2.x)'}\"\n          f\" | peak-norm: {PERCH_PEAK_NORM}\")\n_perch_batched = True\n\ndef _probe_embed_dim():\n    try:\n        return int(_perch_run(np.zeros((1, PERCH_LEN), dtype=\"float32\")).shape[-1])\n    except Exception as ex:\n        print(\"  could not probe embedding dim, assuming 1280:\", type(ex).__name__)\n        return 1280\n\ndef to32k(y):\n    return librosa.resample(np.asarray(y, dtype=\"float32\"), orig_sr=SR, target_sr=PERCH_SR)\n\ndef _peak_norm(batch):\n    \"\"\"DC-remove and scale each window to PERCH_PEAK_NORM -- zoo_interface.normalize_audio.\"\"\"\n    if PERCH_PEAK_NORM is None: return batch\n    b = np.asarray(batch, dtype=\"float32\").copy()\n    b -= np.mean(b, axis=-1, keepdims=True)\n    pk = np.max(np.abs(b), axis=-1, keepdims=True)\n    b = np.divide(b, pk, where=(pk > 0.0))\n    return (b * PERCH_PEAK_NORM).astype(\"float32\")\n\ndef _extract_embedding(out):\n    \"\"\"Perch 2.x -> dict with 'embedding' (pooled) plus 'spatial_embedding'/'frontend'/\n    logit heads. Perch 1.x -> dict with 'embedding', or a bare (logits, embeddings) tuple.\"\"\"\n    if hasattr(out, \"keys\"):\n        if \"embedding\" not in out:\n            raise KeyError(f\"no 'embedding' head in Perch output; got {sorted(out.keys())}\")\n        emb = out[\"embedding\"]\n    else:\n        emb = out[1]\n    return np.asarray(emb.numpy() if hasattr(emb, \"numpy\") else emb)\n\ndef _infer(batch):\n    batch = _peak_norm(batch)\n    if PERCH_BACKEND == \"onnx\":\n        result = _ort_sess.run(None, {_ort_in: batch})\n        out = dict(zip(_ort_out_names, result))   # verified real keys: embedding/\n                                                   # spatial_embedding/spectrogram/label\n    else:\n        out = perch.infer_tf(batch) if _HAS_INFER_TF else _SERVING(inputs=batch)\n    return _extract_embedding(out)\n\ndef _perch_run(waves):\n    global _perch_batched\n    n = len(waves)\n    if n == 0: return np.zeros((0, EMBED_DIM if \"EMBED_DIM\" in globals() else 1280), dtype=\"float32\")\n    if _perch_batched:\n        try:\n            out = []\n            for i in range(0, n, PERCH_BATCH):\n                chunk = waves[i:i+PERCH_BATCH]; k = len(chunk)\n                if k < PERCH_BATCH:\n                    chunk = np.concatenate([chunk, np.zeros((PERCH_BATCH-k, PERCH_LEN), dtype=\"float32\")])\n                out.append(_infer(chunk)[:k])\n            return np.concatenate(out)\n        except Exception as ex:\n            print(\"batched Perch inference failed, falling back to per-clip:\", type(ex).__name__)\n            _perch_batched = False\n    return np.stack([_infer(w[np.newaxis, :])[0] for w in waves])\n\ndef perch_embed(waveforms):\n    \"\"\"Isolated clips -> (n, D). CENTRED zero-padding, Perch's documented convention.\"\"\"\n    ws = []\n    for w in waveforms:\n        w = to32k(w)\n        if len(w) >= PERCH_LEN:\n            w = w[:PERCH_LEN]\n        else:\n            pad = PERCH_LEN - len(w); lo = pad // 2\n            w = np.pad(w, (lo, pad - lo))\n        ws.append(w.astype(\"float32\"))\n    return _perch_run(np.stack(ws)) if ws else np.zeros((0, EMBED_DIM), dtype=\"float32\")\n\ndef perch_embed_ctx(y32, centres_sec):\n    \"\"\"Real 5s windows cut from a 32kHz recording, centred on each event time.\"\"\"\n    ws = []\n    for c in centres_sec:\n        a = int(round(c * PERCH_SR)) - PERCH_LEN // 2\n        a = max(0, min(a, max(0, len(y32) - PERCH_LEN)))\n        w = y32[a:a + PERCH_LEN]\n        if len(w) < PERCH_LEN:\n            pad = PERCH_LEN - len(w); lo = pad // 2\n            w = np.pad(w, (lo, pad - lo))\n        ws.append(w.astype(\"float32\"))\n    return _perch_run(np.stack(ws)) if ws else np.zeros((0, EMBED_DIM), dtype=\"float32\")\n\nEMBED_DIM = _probe_embed_dim()\n_DETECTED = {1280: \"v1\", 1536: \"v2\"}.get(EMBED_DIM, \"unknown\")\nprint(f\"Perch embedding dim: {EMBED_DIM}  (detected {_DETECTED}, requested {PERCH_VERSION})\")\nif _DETECTED != PERCH_VERSION:\n    # Wrong artefact attached is the likeliest cause, and it would otherwise produce\n    # perfectly plausible numbers from the wrong model. Refuse rather than mislead.\n    raise RuntimeError(\n        f\"Perch version mismatch: requested {PERCH_VERSION} but the loaded model emits \"\n        f\"{EMBED_DIM}-d embeddings ({_DETECTED}). Loaded from: {PERCH_HANDLE}\\n\"\n        f\"  -> On Kaggle, attach the model variation matching PERCH_VERSION:\\n\"\n        f\"     v2: bird-vocalization-classifier/tensorFlow2/perch_v2 (GPU, v2) \"\n        f\"or perch_v2_cpu (CPU, v1)\\n\"\n        f\"     v1: bird-vocalization-classifier/tensorFlow2/bird-vocalization-classifier (v8)\")\n\n# ------------------------------------------------- pre-registered target selection\n_counts = meta.primary_label.value_counts()\n_eligible = [sp for sp in _counts.index if _counts[sp] >= MIN_RECS]\nif PRIMARY_TARGET not in _eligible:\n    raise ValueError(f\"{PRIMARY_TARGET} has only {_counts.get(PRIMARY_TARGET, 0)} recordings \"\n                     f\"(need >= {MIN_RECS})\")\nif SINGLE_TARGET_COMPARISON:\n    TARGET_LIST = [PRIMARY_TARGET]\nelse:\n    TARGET_LIST = [PRIMARY_TARGET] + [sp for sp in _eligible if sp != PRIMARY_TARGET][:N_TARGETS - 1]\n    if SMOKE_TEST:\n        TARGET_LIST = TARGET_LIST[:2]\nTARGET_SET = set(TARGET_LIST)     # <-- consumed by CELL 2 to exclude ALL targets from the\n                                  #     encoder bank. v1 hard-coded \"whbsho3\" there, which\n                                  #     silently trained the encoder on every OTHER target.\nprint(f\"\\nPRE-REGISTERED TARGETS (n={len(TARGET_LIST)}, smoke_test={SMOKE_TEST}, \"\n      f\"single_target_comparison={SINGLE_TARGET_COMPARISON}) \"\n      f\"-- copy this list into the logbook BEFORE running:\")\nfor sp in TARGET_LIST:\n    print(f\"   {sp}  ({_counts[sp]} recordings)\")\n\n# ------------------------------------------------- global species cache (ONCE)\n_cand = (meta[~meta.primary_label.isin(TARGET_SET)]['primary_label']\n         .value_counts().head(N_CANDIDATES).index.tolist())\nSPECIES_CACHE = {}\n_t0 = time.time()\nfor sp in _cand:\n    calls = []\n    for f in meta[meta.primary_label == sp]['filename'].tolist()[:CALIB_PER_SPECIES]:\n        try: calls += peak_windows(ld(os.path.join(AUD, f), 30), k=1)\n        except Exception: pass\n    if not calls: continue\n    SPECIES_CACHE[sp] = {\"calls\": calls, \"emb\": l2n(perch_embed(calls)).mean(0)}\nprint(f\"species cache: {len(SPECIES_CACHE)}/{len(_cand)} candidates embedded \"\n      f\"in {time.time()-_t0:.0f}s (computed once, reused by every target)\")\n\n_decoy_pool_cache = {}\ndef _deep_decoy_pool(sp):\n    if sp not in _decoy_pool_cache:\n        pool = []\n        for f in meta[meta.primary_label == sp]['filename'].tolist()[:DECOY_RECORDINGS]:\n            try: pool += peak_windows(ld(os.path.join(AUD, f), 30), k=1)\n            except Exception: pass\n        _decoy_pool_cache[sp] = pool or SPECIES_CACHE[sp][\"calls\"]\n    return _decoy_pool_cache[sp]\n\n# ------------------------------------------------------------------ soundscape beds\nbed_paths = sorted(glob.glob(os.path.join(SS, \"*.ogg\")))[:N_CALIB_BEDS + N_EVAL_BEDS]\nall_beds = [ld(p, 240.0) for p in bed_paths]\ncalib_beds, eval_beds = all_beds[:N_CALIB_BEDS], all_beds[N_CALIB_BEDS:]\nBED_SEC = 240.0\nprint(f\"beds: {len(calib_beds)} calibration + {len(eval_beds)} evaluation (disjoint)\")\n\n# ------------------------------------------------------------ shot selection\ndef choose_shots(files, ratings, embs, n_shot, strategy, rng):\n    \"\"\"-> indices into `files`. Each strategy is a different answer to the open question\n    Nolasco et al. sec.4.1 raise: which 5 examples should a deployed system be given?\n      rated   : highest Xeno-Canto rating (ARGUS v1 behaviour -- best-quality audio)\n      first   : first n in metadata order (DCASE Task 5 behaviour -- no curation)\n      random  : uniform sample (control)\n      diverse : greedy max-min distance in Perch space (maximum coverage of the\n                call repertoire -- the thing the expert annotators said was missing)\n    \"\"\"\n    n = len(files)\n    if strategy == \"rated\":\n        return list(np.argsort(-np.asarray(ratings, dtype=float))[:n_shot])\n    if strategy == \"first\":\n        return list(range(min(n_shot, n)))\n    if strategy == \"random\":\n        return list(rng.permutation(n)[:n_shot])\n    if strategy == \"diverse\":\n        E = l2n(embs)\n        picked = [int(np.argmax(np.asarray(ratings, dtype=float)))]   # deterministic seed point\n        while len(picked) < min(n_shot, n):\n            d = np.min(1.0 - E @ E[picked].T, axis=1)   # cosine distance to nearest picked\n            d[picked] = -np.inf\n            picked.append(int(np.argmax(d)))\n        return picked\n    raise ValueError(f\"unknown shot strategy: {strategy}\")\n\n# ------------------------------------------------------------ stage-2 verifiers\nclass PrototypicalProbe:\n    \"\"\"Prototype-distance classifier. Drop-in for LogisticRegression (fit/predict_proba).\n\n    WHY: Perch 2.0's own benchmark (Table 3, BEANS) reports prototypical probing beating\n    linear probing on DETECTION tasks by +0.078 mAP (0.426 -> 0.504). ARGUS stage 2 is a\n    detection task, and currently uses a linear probe.\n\n    K prototypes per class rather than one mean, following the ProtoPNet lineage Perch 2.0\n    adopts (4 prototypes/class, prediction = max activation across them). This is also the\n    direct answer to Nolasco et al. sec.4.1: a single averaged prototype \"implies that the\n    examples are in some sense all of one kind\" -- false for a species with several call\n    types, which is exactly ARGUS's failure mode.\n\n    Scores are cosine similarities (inputs are L2-normalised upstream), turned into\n    calibrated probabilities by a fitted temperature. Calibration matters here because the\n    fusion rules multiply p*v -- an uncalibrated v that saturates to 0/1 becomes a hard\n    veto and destroys ranking, which is the same trap that forced C=0.1 on the LogReg.\n    \"\"\"\n\n    def __init__(self, n_prototypes=4, random_state=0, class_weight=\"balanced\"):\n        self.n_prototypes = n_prototypes\n        self.random_state = random_state\n        self.class_weight = class_weight\n\n    def _protos(self, Xc):\n        k = int(min(self.n_prototypes, len(Xc)))\n        if k <= 1:\n            return l2n(Xc.mean(0, keepdims=True))\n        km = KMeans(n_clusters=k, n_init=10, random_state=self.random_state).fit(Xc)\n        return l2n(km.cluster_centers_)\n\n    def _scores(self, X):\n        \"\"\"-> (n, 2): max cosine similarity to each class's prototype set.\"\"\"\n        Xn = l2n(X)\n        return np.stack([(Xn @ self.protos_[c].T).max(1) for c in (0, 1)], axis=1)\n\n    @staticmethod\n    def _softmax(z):\n        z = z - z.max(axis=1, keepdims=True)\n        e = np.exp(z)\n        return e / e.sum(axis=1, keepdims=True)\n\n    def _fit_temperature(self, S, y, w, eps=0.05):\n        \"\"\"Temperature by weighted log-loss with LABEL SMOOTHING.\n\n        Smoothing is not cosmetic here. The calibration set separates almost perfectly in\n        Perch space, so an unsmoothed objective drives T -> 0 and saturates predict_proba\n        to hard 0/1. That turns v into a veto rather than a score, which destroys the\n        ranking that AP and the p*v fusion rules depend on -- the identical failure mode\n        that forced C=0.1 on the LogisticRegression verifier. With targets smoothed to\n        1-eps the optimum is finite and the probabilities stay soft.\n        \"\"\"\n        best_T, best_ll = 1.0, np.inf\n        for T in np.geomspace(0.02, 3.0, 60):\n            p_true = np.clip(self._softmax(S / T)[np.arange(len(y)), y], 1e-9, 1.0)\n            p_false = np.clip(1.0 - p_true, 1e-9, 1.0)\n            ll = float(-(w * ((1.0 - eps) * np.log(p_true)\n                              + eps * np.log(p_false))).sum() / w.sum())\n            if ll < best_ll:\n                best_T, best_ll = float(T), ll\n        return best_T\n\n    def fit(self, X, y):\n        X = np.asarray(X, dtype=\"float32\"); y = np.asarray(y).astype(int)\n        self.protos_ = {c: self._protos(X[y == c]) for c in (0, 1)}\n        if self.class_weight == \"balanced\":\n            cw = {c: len(y) / (2.0 * max(1, (y == c).sum())) for c in (0, 1)}\n        else:\n            cw = {0: 1.0, 1: 1.0}\n        w = np.array([cw[int(v)] for v in y], dtype=\"float64\")\n        self.temperature_ = self._fit_temperature(self._scores(X), y, w)\n        return self\n\n    def predict_proba(self, X):\n        return self._softmax(self._scores(X) / self.temperature_)\n\n    def score(self, X, y):\n        return float((self.predict_proba(X).argmax(1) == np.asarray(y).astype(int)).mean())\n\n\ndef _make_verifier():\n    if VERIFIER == \"logreg\":\n        # C deliberately small: 1280/1536-d embeddings with ~100 samples separate perfectly\n        # at C=1, saturating predict_proba to 0/1 and turning v into a hard veto.\n        return LogisticRegression(C=0.1, max_iter=2000, class_weight=\"balanced\")\n    if VERIFIER == \"proto\":\n        return PrototypicalProbe(n_prototypes=4, random_state=0)\n    raise ValueError(f\"unknown VERIFIER: {VERIFIER}\")\n\n\n# ------------------------------------------------------------ per-species covariates\ndef _covariates(support_views, own_embs, decoy_sims):\n    \"\"\"Recorded so the difficulty-factor regression Nolasco et al. could not resolve\n    becomes possible later. Free now; unreconstructable after the fact.\"\"\"\n    E = l2n(own_embs)\n    if len(E) > 1:\n        S = E @ E.T\n        iu = np.triu_indices(len(E), k=1)\n        stereotypy = float(S[iu].mean())        # high = calls resemble each other\n    else:\n        stereotypy = float('nan')\n    cents, bws, durs = [], [], []\n    for w in support_views:\n        cents.append(float(np.mean(librosa.feature.spectral_centroid(y=w, sr=SR))))\n        bws.append(float(np.mean(librosa.feature.spectral_bandwidth(y=w, sr=SR))))\n        e = np.square(w, dtype=np.float64)\n        thr = 0.1 * e.max() if e.max() > 0 else 0.0\n        durs.append(float((e > thr).sum()) / SR)   # seconds above 10% of peak energy\n    return dict(\n        stereotypy=stereotypy,\n        spectral_centroid_hz=float(np.mean(cents)),\n        spectral_bandwidth_hz=float(np.mean(bws)),\n        call_duration_s=float(np.mean(durs)),\n        decoy_sim_top1=float(max(decoy_sims)) if decoy_sims else float('nan'),\n        decoy_sim_mean=float(np.mean(decoy_sims)) if decoy_sims else float('nan'),\n    )\n\n# ------------------------------------------------------------ per-target builder\n# Everything except the verifier fit is deterministic given (species, n_shot, strategy,\n# seed) -- and it is nearly all Perch inference, the dominant cost. The ablation ladder\n# rebuilds targets once per rung, so without this cache every rung would re-embed the same\n# audio. Cached across rungs; only the (cheap) classifier is refitted.\n_TARGET_CACHE = {}\n\ndef build_target(sp, n_shot=N_SHOT, strategy=SHOT_STRATEGY, seed=0, verbose=True,\n                  decoy_mode=\"hard\", recordist_disjoint=False):\n    \"\"\"Everything v1 did at module level for one hard-coded species, now per species.\n    Returns a self-contained context dict consumed by CELL 2.\n\n    decoy_mode selects WHICH species (out of every candidate ranked by Perch-embedding\n    similarity to this target) become the N_DECOYS used at test time -- roadmap 3.3:\n      'hard'   (default) -- the top N_DECOYS nearest. Unchanged behaviour; every result\n                published before this flag existed used this mode.\n      'medium' -- ranks N_DECOYS..2*N_DECOYS-1 (5th-8th nearest at N_DECOYS=4).\n      'random' -- N_DECOYS drawn uniformly from the full non-target candidate pool,\n                  seeded from this call's own `seed` so it's reproducible, not a\n                  fresh draw on every call.\n    Only what a run is SCORED against changes; the verifier is still trained against\n    every non-decoy candidate as negatives regardless of mode (calib_neg_species below),\n    so 'easier' decoys are a harder TEST, not a easier training signal.\n\n    recordist_disjoint -- confound-closing run 3 (roadmap \"owed\" list): \"did it learn the\n    microphone, not the bird?\" Default False reproduces every previously-published number\n    byte-for-byte (held-out pool = every recording not used as a shot, regardless of who\n    recorded it). True additionally excludes, from the held-out pool, any recording whose\n    RECORDIST_COL matches one of the n_shot support recordings -- so if a species'\n    apparent AUC were actually partly \"recognise this recordist's mic/room\", a\n    recordist-disjoint held-out pool would show it as a drop. Can raise ValueError if a\n    species has too few distinct recordists (the held-out pool empties out) -- that itself\n    is worth logging, not a bug to route around.\"\"\"\n    # EMBED_BACKEND is part of the key for the same reason recordist_disjoint is: the cache\n    # stores the fitted verifier's TRAINING MATRIX (c[\"X\"]) and refits only the cheap\n    # classifier on it. Without the backend in the key, switching to the distilled encoder\n    # would silently refit on Perch's cached embeddings and report the result as distilled.\n    ck = (sp, int(n_shot), strategy, int(seed), decoy_mode, bool(recordist_disjoint),\n          EMBED_BACKEND)\n    if ck in _TARGET_CACHE:\n        c = _TARGET_CACHE[ck]\n        clf = _make_verifier().fit(c[\"X\"], c[\"yv\"])\n        cov = dict(c[\"cov\"]); cov[\"clf_train_acc\"] = float(clf.score(c[\"X\"], c[\"yv\"]))\n        cov[\"verifier\"] = VERIFIER\n        if verbose:\n            print(f\"  [{sp}] cached build reused; verifier={VERIFIER} \"\n                  f\"(train acc {cov['clf_train_acc']:.2f})\")\n        return dict(species=sp, support=c[\"support\"], target_pool=c[\"target_pool\"],\n                    PROTO=c[\"PROTO\"], DECOYS=c[\"DECOYS\"], decoy_pools=c[\"decoy_pools\"],\n                    perch_clf=clf, covariates=cov, neg_calls=c[\"neg_calls\"])\n    rng = np.random.default_rng(seed)\n    rows = meta[meta.primary_label == sp]\n    files   = rows['filename'].tolist()\n    ratings = rows['rating'].tolist() if 'rating' in rows else [0.0] * len(files)\n    if recordist_disjoint and RECORDIST_COL not in rows:\n        raise KeyError(f\"recordist_disjoint=True but {RECORDIST_COL!r} is not a column \"\n                        f\"of train_metadata.csv; have {rows.columns.tolist()}\")\n    authors = (rows[RECORDIST_COL].tolist() if recordist_disjoint\n               else [None] * len(files))\n    if len(files) < n_shot + 2:\n        raise ValueError(f\"not enough recordings for {sp}: {len(files)}\")\n\n    raw = [ld(os.path.join(AUD, f), 30) for f in files]\n    # 'diverse' needs one embedding per recording to pick a spread; others don't.\n    rec_embs = (perch_embed([peak_windows(y, k=1)[0] for y in raw])\n                if strategy == \"diverse\" else np.zeros((len(raw), 1)))\n\n    shot_idx = choose_shots(files, ratings, rec_embs, n_shot, strategy, rng)\n    shot_set = set(int(i) for i in shot_idx)\n\n    support = build_support_views([raw[i] for i in shot_idx])\n    # Held-out pool = every recording NOT used as a shot. v1 sliced files[N_SHOT:],\n    # which under any non-'first' strategy would have leaked support into evaluation.\n    excluded_idx = set(shot_set)\n    if recordist_disjoint:\n        # pandas gives every missing value in a string column the SAME NaN singleton, and\n        # `a in shot_authors` does an identity check before equality -- so a NaN-authored\n        # shot correctly excludes every other NaN-authored recording as one shared\n        # pseudo-recordist. This is conservative (over-exclusion, shrinking the held-out\n        # pool), not a leak -- verified 28 Aug 2026 against the real pandas behaviour after\n        # an audit initially suspected the opposite (NaN != NaN silently under-excluding).\n        # It fails safe, so the only thing worth doing is making it visible.\n        n_nan = sum(1 for a in authors if a is not None and a != a)\n        if n_nan and verbose:\n            print(f\"  [{sp}] {n_nan}/{len(authors)} recordings have a missing author -- \"\n                  f\"treated as one shared pseudo-recordist for recordist_disjoint\")\n        shot_authors = {authors[i] for i in shot_set if authors[i] is not None}\n        excluded_idx |= {i for i, a in enumerate(authors) if a in shot_authors}\n    target_pool = [w for i, y in enumerate(raw) if i not in excluded_idx\n                   for w in peak_windows(y, k=1)]\n    if not target_pool:\n        extra = (f\" ({len(excluded_idx) - len(shot_set)} extra recordings excluded for \"\n                 f\"sharing a recordist with a shot)\" if recordist_disjoint else \"\")\n        raise ValueError(f\"{sp}: no held-out recordings left after taking {n_shot} \"\n                          f\"shots{extra}\")\n\n    own_embs = perch_embed(support)\n    PROTO = l2n(own_embs).mean(0); PROTO /= np.linalg.norm(PROTO) + 1e-9\n\n    pool_sp = {k: v for k, v in SPECIES_CACHE.items() if k != sp}\n    ranked  = sorted(pool_sp, key=lambda o: -cos(pool_sp[o][\"emb\"], PROTO))\n    if decoy_mode == \"hard\":\n        DECOYS = ranked[:N_DECOYS]\n    elif decoy_mode == \"medium\":\n        DECOYS = ranked[N_DECOYS:2 * N_DECOYS]\n        if len(DECOYS) < N_DECOYS:\n            raise ValueError(f\"{sp}: only {len(ranked)} candidates ranked, need \"\n                              f\"{2*N_DECOYS} for decoy_mode='medium'\")\n    elif decoy_mode == \"random\":\n        # Seeded from this call's `seed`, offset so it never coincides with any other\n        # rng drawn from the same seed in this function (shot choice, calibration).\n        DECOYS = list(np.random.default_rng(90_000 + seed).choice(\n            ranked, size=N_DECOYS, replace=False))\n    else:\n        raise ValueError(f\"unknown decoy_mode: {decoy_mode!r} (want hard|medium|random)\")\n    decoy_sims  = [cos(pool_sp[o][\"emb\"], PROTO) for o in DECOYS]\n    decoy_pools = {o: _deep_decoy_pool(o) for o in DECOYS}\n    calib_neg_species = [o for o in ranked if o not in DECOYS]\n\n    # ---- domain-matched perch_clf (identical protocol to v1, per target) ----\n    neg_calls = [w for o in calib_neg_species for w in pool_sp[o][\"calls\"]]\n    rng_cal = np.random.default_rng(11 + seed)\n    X_parts, y_all = [], []\n    for bed in calib_beds:\n        y_bed = bed.copy(); spans, centres, labels = [], [], []\n        plan = [(1, support)] * 12 + [(0, neg_calls)] * 12\n        plan = [plan[i] for i in rng_cal.permutation(len(plan))]\n        for i, (lab, pl) in enumerate(plan):\n            c = place(y_bed, pl, rng_cal, spans, SNR_LEVELS[i % len(SNR_LEVELS)])\n            if c is None: continue\n            centres.append(c + float(rng_cal.uniform(-0.25, 0.25))); labels.append(lab)\n        for _ in range(8):\n            pos = int(rng_cal.integers(0, len(y_bed) - CLEN))\n            if any(pos < e and pos + CLEN > s for s, e in spans): continue\n            centres.append((pos + CLEN/2) / SR); labels.append(0)\n        X_parts.append(verifier_embed_ctx(y_bed, centres)); y_all += labels\n\n    X = l2n(np.concatenate(X_parts)); yv = np.array(y_all)\n    clf = _make_verifier().fit(X, yv)\n\n    cov = _covariates(support, own_embs, decoy_sims)\n    cov.update(species=sp, n_recordings=len(files), n_shot=n_shot,\n               shot_strategy=strategy, n_heldout_calls=len(target_pool),\n               clf_train_acc=float(clf.score(X, yv)),\n               decoy_mode=decoy_mode,\n               recordist_disjoint=bool(recordist_disjoint),\n               n_distinct_recordists=(len(set(authors)) if recordist_disjoint else -1),\n               # provenance: a results CSV that does not say which Perch produced it\n               # is not comparable to any other run.\n               perch_version=PERCH_VERSION, perch_embed_dim=EMBED_DIM,\n               perch_peak_norm=(PERCH_PEAK_NORM if PERCH_PEAK_NORM is not None else 0.0),\n               perch_backend=PERCH_BACKEND,\n               verifier=VERIFIER)\n\n    if verbose:\n        rd_note = (f\"  distinct_recordists={cov['n_distinct_recordists']}\"\n                   if recordist_disjoint else \"\")\n        print(f\"  [{sp}] shots={strategy}({n_shot})  heldout={len(target_pool)} calls  \"\n              f\"decoys={DECOYS} ({decoy_mode})  top-sim={cov['decoy_sim_top1']:.3f}  \"\n              f\"stereotypy={cov['stereotypy']:.3f}  verifier={VERIFIER}{rd_note}\")\n\n    _TARGET_CACHE[ck] = dict(support=support, target_pool=target_pool, PROTO=PROTO,\n                             DECOYS=DECOYS, decoy_pools=decoy_pools, X=X, yv=yv,\n                             neg_calls=neg_calls,\n                             cov={k: v for k, v in cov.items() if k != \"clf_train_acc\"})\n    return dict(species=sp, support=support, target_pool=target_pool, PROTO=PROTO,\n                DECOYS=DECOYS, decoy_pools=decoy_pools, perch_clf=clf, covariates=cov,\n                neg_calls=neg_calls)\n\nprint(\"\\nCELL 1 v2 ready. build_target(sp) is the entry point; TARGET_SET is exported \"\n      \"for the encoder-bank exclusion in CELL 2.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"cbd8d13a","cell_type":"markdown","source":"## Cell 2 — Two-Stage Detection, Multi-Species\n\n`arguswala_v2.py`","metadata":{}},{"id":"227ac64c","cell_type":"code","source":"# ===== CELL 2 v2 / TWO-STAGE across N species =====\n# Needs CELL 1 v2 (arguswala_setup_v2.py).\n#\n# WHAT CHANGED vs v1:\n#\n#  (1) THE BUG. v1 line 216 read:\n#          sps = [s for s in meta.primary_label.unique() if s != \"whbsho3\"]\n#      The target was hard-coded while CELL 1 defined it as a variable. The moment a\n#      second target is evaluated, that line puts the new target INTO the encoder's\n#      training bank -- the encoder is pre-trained on the species it must never have\n#      seen, every number inflates, and nothing crashes. Now excludes all of TARGET_SET,\n#      with an assertion so it can never regress silently.\n#\n#  (2) ONE ENCODER, MANY TARGETS. Legitimate because no target is in the bank, and it\n#      turns O(species x encoder-trainings) into O(1).\n#\n#  (3) PURE-NEGATIVE CONTROL (neg_pass). Nothing in v1 ever ran on a recording with zero\n#      target calls -- in deployment that is the common case, and false alarms per hour\n#      on clean audio is the first number a conservation biologist asks for.\n#\n#  (4) FEATURE TOGGLE. Liang et al. (DCASE 2024 baseline) Table III: PCEN gives the best\n#      PRECISION of any front-end tested (68.0%) while log-Mel gives the best F1 (63.7 vs\n#      60.0). v1 assumed PCEN. Now it is a measured choice.\n#\n#  (5) COVARIATES + per-species results written to CSV, so the difficulty-factor analysis\n#      Nolasco et al. attempted and failed can actually be run afterwards.\nimport numpy as np, os, glob, copy, time, librosa, torch, torch.nn as nn, torch.optim as optim\nimport pandas as pd\n\nSR, N_FFT, HOP, N_MELS = 22050, 1024, 256, 128   # identical to the DCASE 2024 baseline\n                                                 # (22.05kHz / 1024 / 256 / 128) -- keep it\n                                                 # that way, it is what makes a DCASE run cheap\nSEG_SEC = 1.0; SEG_FR = int(np.ceil(SEG_SEC*SR/HOP)); CLEN = int(SEG_SEC*SR); EMB = 128\nFT_STEPS, FT_LR, FT_BATCH = 250, 1e-4, 16\nN_INJECT, N_NEG, INJ_SNR = 45, 200, (-6, 12)\nCAND_FLOOR, STRIDE, IOU_THR = 0.10, max(1, SEG_FR//3), 0.30\nMAX_EVENT_SEC = 10.0\nCALLS_PER_BED = 12\n\nFEATURE = \"pcen\"          # 'pcen' | 'logmel'\nSEEDS   = (7, 8, 9)       # raise to >=10 once the timing measurement says it fits\nENCODER_SEEDS = (0, 1, 2) # >1 entry measures encoder-training variance (roadmap 3.4).\n                          # Widened 14 Aug 2026 after the first real run: whbsho3 stage-1\n                          # AUC came back 0.517 here vs. the earlier single-target baseline\n                          # of 0.459. This tests ONE candidate cause -- training-run\n                          # stochasticity on a FIXED bank (build_banks() always uses\n                          # seed=0 regardless of this tuple, so bank COMPOSITION doesn't\n                          # vary here). A separate single-target comparison run is still\n                          # the only way to test the other candidate cause (bank\n                          # composition differs because this run excludes 2 species from\n                          # the encoder bank, the original baseline only excluded 1).\n\n# 18 Aug 2026: a Kaggle GPU-quota cutoff killed a run mid-way through CELL 4 after\n# CELL 2's 5.5-hour species loop had already completed and its CSV was safely\n# downloaded. Restarting the whole notebook to reach CELL 4 again would re-pay that\n# 5.5 hours for a result already in hand. SKIP_MAIN_LOOP=True runs bank-building and\n# encoder training (~10-20 min, everything CELL 4 actually needs) and stops there --\n# `res`/`wide`/the printed report below never run, and never need to. CELL 4's own\n# checkpointing (RESUME=True) then picks up the ladder at the first unfinished config.\n# Leave False for a normal full run.\nSKIP_MAIN_LOOP = True\n\n# ---------------------------------------------------------------- ablation ladder\n# Stage-1 rungs. Both default OFF so rung 0 reproduces the original pipeline exactly.\n# (Stage-2 rung lives in CELL 1 as VERIFIER = 'logreg' | 'proto'.)\nUSE_TIMEFILTER = False    # rung 2: TimeFilterAug on the support during fine-tuning\nUSE_HARD_NEG   = False    # rung 3: hard-negative mining in the fine-tuning loop\nHARD_NEG_KEEP  = 0.25     # fraction of candidate negatives retained as \"hard\"\n\n# rung 5 (CELL 5, confound-closing run 2 -- roadmap \"owed\" list): stage-1's per-target\n# fine-tuning has always drawn its negatives as arbitrary random crops of the bed\n# (whatever silence/wind/distant noise happens to be there) -- never a labelled OTHER\n# SPECIES call. The stage-2 Perch verifier already trains against real other-species\n# calls (calib_neg_species in build_target); stage 1 never has. Default OFF reproduces\n# every previously-published number byte-for-byte. When True, stage1() injects real\n# calls from ctx['neg_calls'] (the SAME non-decoy species pool the verifier uses) into a\n# copy of the clean recording and crops THOSE positions as negatives -- but N_NEG=200\n# 1s calls with 0.4s gaps need >=280s in a 240s bed, which is geometrically impossible\n# (measured: rejection-sampling places ~126 of the 200 target, min 122/max 130 over 10\n# draws), so on THIS side the \"instead of arbitrary background\" framing overclaims --\n# ~74 of the ~200 negatives are still background top-up on every call (see the top-up\n# branch in stage1() below). What this rung actually tests is \"~126 species-matched +\n# ~74 background\" vs \"~200 background\", not a clean substitution. Restated honestly in\n# both places 28 Aug 2026, after an audit caught the original wording.\nUSE_SPECIES_NEG = False\n\n# CELL 5, confound-closing run 1: how far a sliding Perch-embedding + verifier scan\n# steps between windows when scoring a recording with NO stage-1 CNN in the loop at all\n# (detect_perch_only below). Perch inference is the pipeline's dominant per-call cost, so\n# this trades runtime against localisation precision; narrow it only after the first\n# timing print says the wider stride is affordable.\nPERCH_ALONE_STRIDE_S = 0.5\n\ndef timefilteraug(x, rng, m_range=(3, 6), lo_db=-6.0, hi_db=8.0):\n    \"\"\"TimeFilterAug -- Zou et al. 2024 (DCASE 2023 Task 5, 1st place, 63.8% F), eq. 6.\n\n    Motivation, in their words: the five shots \"are typically clear, whereas the query set\n    for predictions often contains interference noise, predominantly from far-field sound\n    and background impulse noise.\" This simulates that by applying a random piecewise-\n    LINEAR gain envelope across TIME (not a mask, and not across frequency):\n\n      partition the window into m in [3,6] segments; draw a breakpoint gain g_i ~ U[0,1]\n      per boundary; within each segment interpolate linearly between adjacent g's; map\n      that ramp onto [lo_db, hi_db] = [-6, +8] dB; apply.\n\n    Their front-end is PCEN at 22050 Hz / n_fft 1024 / hop 256 -- identical to ARGUS, so\n    the dB bounds transfer directly rather than needing rescaling.\n    \"\"\"\n    T = x.shape[1]\n    m = int(rng.integers(m_range[0], m_range[1] + 1))\n    if T < m + 1:\n        return x\n    cuts  = np.sort(rng.choice(np.arange(1, T), size=m - 1, replace=False))\n    edges = np.concatenate([[0], cuts, [T]])\n    g = rng.random(m + 1)\n    alpha = np.concatenate([\n        np.linspace(g[i], g[i + 1], int(edges[i + 1] - edges[i]), endpoint=False)\n        for i in range(m)\n    ])\n    beta = (lo_db + (hi_db - lo_db) * alpha).astype(\"float32\")   # per-frame gain in dB\n    if FEATURE == \"logmel\":\n        # amplitude_to_db output is ALREADY logarithmic -> a dB gain is additive.\n        return x + beta[np.newaxis, :]\n    # PCEN is a linear-domain magnitude -> a dB gain is multiplicative.\n    return x * (10.0 ** (beta / 20.0))[np.newaxis, :].astype(\"float32\")\n\ndef aug_pos(x, rng):\n    \"\"\"Augmentation applied to positives (the injected support) during fine-tuning.\"\"\"\n    x = specaug(x, rng)\n    if USE_TIMEFILTER:\n        x = timefilteraug(x, rng)\n    return x\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\"); print(\"device:\", device)\n\ndef rms(x): return float(np.sqrt(np.mean(np.square(x, dtype=np.float64)) + 1e-12))\ndef ld(p, dur=None):\n    y, _ = librosa.load(p, sr=SR, mono=True, duration=dur); return y.astype(\"float32\")\n\nFEATURES = (\"pcen\", \"logmel\")   # every front-end an encoder is trained for\n\ndef feat(y, feature=None):\n    \"\"\"Front-end. PCEN vs log-Mel is an experiment, not an assumption.\n\n    `feature=None` reads the FEATURE global -- that is what makes the ladder able to\n    switch front-ends by rebinding a global. Pass it explicitly when featurising for a\n    front-end other than the currently-selected one (e.g. building both encoder banks).\n    \"\"\"\n    feature = feature or FEATURE\n    m = librosa.feature.melspectrogram(y=y, sr=SR, n_fft=N_FFT, hop_length=HOP,\n                                       n_mels=N_MELS, power=1.0)\n    if feature == \"pcen\":\n        return librosa.pcen(m * (2**31), sr=SR, hop_length=HOP).astype(\"float32\")\n    if feature == \"logmel\":\n        return librosa.amplitude_to_db(m, ref=np.max).astype(\"float32\")\n    raise ValueError(f\"unknown feature: {feature}\")\n\ndef crop(f, c):\n    a = c - SEG_FR//2; p = np.full((N_MELS, SEG_FR), f.min(), dtype=\"float32\")\n    lo, hi = max(0, a), min(f.shape[1], a+SEG_FR)\n    if hi > lo: p[:, (lo-a):(lo-a)+(hi-lo)] = f[:, lo:hi]\n    return p\ndef loud_off(y, n=CLEN):\n    if len(y) <= n: return 0\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    return int(np.argmax(cs[n:]-cs[:-n]))\ndef specaug(x, rng):\n    if rng.random() >= 0.6: return x\n    x = x.copy(); fl = x.min()\n    for _ in range(2):\n        k = int(rng.integers(0, 21))\n        if k: a = int(rng.integers(0, max(1, N_MELS-k))); x[a:a+k, :] = fl\n        k = int(rng.integers(0, 11))\n        if k: a = int(rng.integers(0, max(1, SEG_FR-k))); x[:, a:a+k] = fl\n    return x\ndef inject(bed, calls, n, snr, rng, gap_s=0.4):\n    y = bed.astype(\"float32\").copy(); L = len(y); gap = int(gap_s*SR); sp = []\n    if L <= CLEN: return y, sp\n    for _ in range(n*60):\n        if len(sp) >= n: break\n        pos = int(rng.integers(0, L-CLEN))\n        if any(pos < e+gap and pos+CLEN > s-gap for s, e in sp): continue\n        c = calls[int(rng.integers(len(calls)))][:CLEN]\n        if len(c) < CLEN: c = np.pad(c, (0, CLEN-len(c)))\n        g = rms(y[pos:pos+CLEN])*(10**(float(rng.uniform(*snr))/20))/(rms(c)+1e-9)\n        y[pos:pos+CLEN] += c.astype(\"float32\")*g; sp.append((pos, pos+CLEN))\n    return y, sorted(sp)\ndef iou(a, b):\n    it = max(0.0, min(a[1], b[1]) - max(a[0], b[0]))\n    u = (a[1]-a[0]) + (b[1]-b[0]) - it\n    return it/u if u > 0 else 0.0\n\nclass Enc(nn.Module):\n    def __init__(s, emb=EMB):\n        super().__init__()\n        def B(i, o): return nn.Sequential(nn.Conv2d(i, o, 3, padding=1, bias=False),\n                                          nn.BatchNorm2d(o), nn.ReLU(), nn.MaxPool2d(2))\n        s.e = nn.Sequential(B(1,128), B(128,128), B(128,128), B(128,emb)); s.p = nn.AdaptiveAvgPool2d(1)\n    def forward(s, x): return s.p(s.e(x.unsqueeze(1))).flatten(1)\n\ndef train_enc(bank, rng, episodes=1000, nway=6, k=4, q=4):\n    m = Enc().to(device); opt = optim.Adam(m.parameters(), 1e-3); ce = nn.CrossEntropyLoss()\n    cl = list(bank); m.train()\n    for ep in range(episodes):\n        ch = rng.choice(cl, size=min(nway, len(cl)), replace=False)\n        sx, sy, qx, qy = [], [], [], []\n        for lab, c in enumerate(ch):\n            idx = rng.permutation(len(bank[c]))[:k+q]\n            for j, i in enumerate(idx):\n                (sx if j < k else qx).append(torch.tensor(specaug(bank[c][i], rng)))\n                (sy if j < k else qy).append(lab)\n        sx = torch.stack(sx).to(device); qx = torch.stack(qx).to(device)\n        sy = torch.tensor(sy).to(device); qy = torch.tensor(qy).to(device)\n        se, qe = m(sx), m(qx)\n        pr = torch.stack([se[sy == c].mean(0) for c in torch.unique(sy)])\n        loss = ce(-torch.cdist(qe, pr), qy); opt.zero_grad(); loss.backward(); opt.step()\n        if (ep+1) % 250 == 0: print(f\"    enc ep {ep+1}/{episodes} loss {loss.item():.3f}\")\n    return m\n\ndef stitch(frames, probs):\n    ab = probs >= CAND_FLOOR; ev = []; i = 0\n    while i < len(frames):\n        if not ab[i]: i += 1; continue\n        j = i\n        while j+1 < len(frames) and ab[j+1] and frames[j+1]-frames[j] <= STRIDE*1.5: j += 1\n        pk = i + int(np.argmax(probs[i:j+1]))\n        s = max(0.0, frames[i]*HOP/SR - SEG_SEC/2); e = frames[j]*HOP/SR + SEG_SEC/2\n        if e - s > MAX_EVENT_SEC:\n            c = frames[pk]*HOP/SR; s, e = max(0.0, c-SEG_SEC/2), c+SEG_SEC/2\n        ev.append((s, e, float(probs[pk]))); i = j+1\n    return ev\n\ndef stitch_time(times, probs, floor=CAND_FLOOR, gap_s=None, max_event_sec=MAX_EVENT_SEC,\n                 seg_sec=SEG_SEC):\n    \"\"\"Same merge-adjacent-above-floor logic as stitch(), but for detectors that never\n    build a mel spectrogram (CELL 5's Perch-alone-as-detector), so positions are already\n    in seconds rather than STRIDE-spaced spectrogram frame indices. `times` must be sorted\n    ascending. gap_s defaults to a spacing consistent with stitch()'s STRIDE*1.5 in frames,\n    converted to seconds.\"\"\"\n    gap_s = gap_s if gap_s is not None else (STRIDE * HOP / SR) * 1.5\n    times = np.asarray(times, dtype=\"float64\")\n    ab = np.asarray(probs) >= floor; ev = []; i = 0\n    while i < len(times):\n        if not ab[i]: i += 1; continue\n        j = i\n        while j+1 < len(times) and ab[j+1] and times[j+1]-times[j] <= gap_s: j += 1\n        pk = i + int(np.argmax(probs[i:j+1]))\n        s = max(0.0, times[i] - seg_sec/2); e = times[j] + seg_sec/2\n        if e - s > max_event_sec:\n            c = times[pk]; s, e = max(0.0, c - seg_sec/2), c + seg_sec/2\n        ev.append((s, e, float(probs[pk]))); i = j+1\n    return ev\n\ndef stage1(enc, rec, sup, rng, neg_species_calls=None):\n    n_inj = max(8, int(round(N_INJECT * len(rec) / (240.0 * SR))))\n    aw, sp = inject(rec, sup, n_inj, INJ_SNR, rng)\n    af, sf = feat(aw), feat(rec)\n    pc = [((a+b)//2)//HOP for a, b in sp]\n    if not pc: return []\n    pos = [crop(af, c) for c in pc]\n    g = SEG_FR; hi = af.shape[1]-g\n\n    def _background_neg():\n        out = []\n        for _ in range(N_NEG*40):\n            if len(out) >= N_NEG: break\n            f = int(rng.integers(g, max(g+1, hi)))\n            if all(abs(f-p) > SEG_FR for p in pc): out.append(crop(af, f))\n        return out\n\n    if USE_SPECIES_NEG and neg_species_calls:\n        # Inject real OTHER-SPECIES calls into a FRESH copy of the clean recording (not\n        # `aw`, which already carries the positive injections -- a separate copy means\n        # positive and negative spans can never collide or corrupt each other) and crop\n        # those exact centres as negatives. N_NEG=200 1s calls at 0.4s gaps cannot all\n        # fit in a 240s bed (needs >=280s) -- inject() places ~126 of them, and the\n        # len(neg)<N_NEG branch below tops up the rest with background. This arm is\n        # therefore \"~126 species-matched + ~74 background\", not a clean 200-call\n        # substitution -- see USE_SPECIES_NEG's definition above for the real numbers.\n        nw, nsp = inject(rec, neg_species_calls, N_NEG, INJ_SNR, rng)\n        nf = feat(nw)\n        nc = [((a+b)//2)//HOP for a, b in nsp]\n        neg = [crop(nf, c) for c in nc if g <= c <= nf.shape[1]-g]\n        if len(neg) < N_NEG:\n            # short recording / crowded bed couldn't place enough species-matched calls --\n            # top up with background rather than silently training on fewer negatives.\n            neg += _background_neg()\n    else:\n        neg = _background_neg()\n    if not neg:\n        neg = [np.full((N_MELS, SEG_FR), af.min(), dtype=\"float32\")]\n    m = copy.deepcopy(enc).to(device); h = nn.Linear(EMB, 2).to(device)\n    opt = optim.Adam(list(m.parameters())+list(h.parameters()), FT_LR)\n    ce = nn.CrossEntropyLoss(); m.train(); h.train(); hf = FT_BATCH//2\n\n    def _train(steps, neg_pool):\n        for _ in range(steps):\n            xs  = [aug_pos(pos[int(rng.integers(len(pos)))], rng) for _ in range(hf)]\n            xs += [neg_pool[int(rng.integers(len(neg_pool)))] for _ in range(hf)]\n            xb = torch.tensor(np.stack(xs)).to(device)\n            yb = torch.tensor([1]*hf+[0]*hf).to(device)\n            loss = ce(h(m(xb)), yb); opt.zero_grad(); loss.backward(); opt.step()\n\n    def _score(items):\n        m.eval(); h.eval(); out = []\n        with torch.no_grad():\n            for i in range(0, len(items), 256):\n                b = torch.tensor(np.stack(items[i:i+256])).to(device)\n                out.append(torch.softmax(h(m(b)), 1)[:, 1].cpu().numpy())\n        m.train(); h.train()\n        return np.concatenate(out)\n\n    if USE_HARD_NEG and len(neg) >= 8:\n        # Negative choice is named as decisive twice in the literature: Nolasco et al.\n        # sec.4.2 (\"prototype-based meta-learning works well when taking care about ...\n        # the choice of negative examples\") and Liang et al., who measure +5.31 F1 from\n        # negative hard sampling. Random background windows are mostly trivial negatives\n        # -- silence and wind -- so half the budget is spent learning nothing. Train\n        # briefly, find the background the model currently CONFUSES with the target, and\n        # spend the rest of the budget there.\n        half = FT_STEPS // 2\n        _train(half, neg)\n        k = max(4, int(round(len(neg) * HARD_NEG_KEEP)))\n        hard = [neg[i] for i in np.argsort(-_score(neg))[:k]]\n        _train(FT_STEPS - half, hard)\n    else:\n        _train(FT_STEPS, neg)\n    m.eval(); h.eval()\n    fr = list(range(g, max(g+1, sf.shape[1]-g), STRIDE)); pb = np.empty(len(fr), dtype=\"float32\")\n    with torch.no_grad():\n        for i in range(0, len(fr), 256):\n            b = torch.tensor(np.stack([crop(sf, f) for f in fr[i:i+256]])).to(device)\n            p = torch.softmax(h(m(b)), 1)[:, 1]; pb[i:i+len(p)] = p.cpu().numpy()\n    return stitch(fr, pb)\n\ndef detect_pv(enc, ctx, rec, rng):\n    \"\"\"-> [(start, end, p, v)]. Target context is now explicit, not a global.\"\"\"\n    ev = stage1(enc, rec, ctx[\"support\"], rng, neg_species_calls=ctx.get(\"neg_calls\"))\n    if not ev: return []\n    centres = [(s + e) / 2 for s, e, _ in ev]\n    # verifier_embed_ctx (CELL 1) dispatches on EMBED_BACKEND. build_target() FITS\n    # ctx[\"perch_clf\"] through the same function, so the classifier is always scored in the\n    # space it was fit in -- under EMBED_BACKEND=\"distilled\" no Perch inference runs here.\n    emb = l2n(verifier_embed_ctx(rec, centres))\n    v = ctx[\"perch_clf\"].predict_proba(emb)[:, 1]\n    return [(s, e, p, float(vi)) for (s, e, p), vi in zip(ev, v)]\n\ndef detect_perch_only(enc, ctx, rec, rng, stride_s=None):\n    \"\"\"-> [(start, end, p, v)], p := v. CELL 5, confound-closing run 1: \"what did the CNN\n    contribute?\" No stage-1 encoder, no mel front-end, no fine-tuning loop anywhere in\n    this path -- Perch embeddings are computed directly on a sliding window across the\n    raw recording and scored with the SAME calibrated verifier (ctx['perch_clf']) every\n    other detector in this file uses, then merged into events with stitch_time(). `enc`\n    and `rng` are accepted only for call-site parity with detect_pv (so eval_pass /\n    spec_pass / neg_pass can swap detectors via one `detect_fn` parameter) -- this path\n    is deterministic given `rec`, so both are unused. p is set equal to v so the FUSIONS\n    machinery still works mechanically, but only the \"perch only (v)\" fusion is a\n    meaningful score here -- any p-dependent fusion rule applied to this detector's\n    output is measuring nothing, since p carries no independent information.\"\"\"\n    stride_s = stride_s if stride_s is not None else PERCH_ALONE_STRIDE_S\n    L = len(rec) / SR\n    times = ([L / 2.0] if L <= SEG_SEC else\n             list(np.arange(SEG_SEC / 2.0, L - SEG_SEC / 2.0 + 1e-9, stride_s)))\n    if not times: return []\n    emb = l2n(perch_embed_ctx(to32k(rec), times))\n    v = ctx[\"perch_clf\"].predict_proba(emb)[:, 1]\n    # max_event_sec=SEG_SEC and an explicit gap_s are BOTH load-bearing; without them this\n    # detector is at serious risk of scoring zero true positives in its normal operating\n    # range, and failing silently in the direction of our own pre-registered prediction.\n    # Geometry: Perch's window is 5 s, so a 1 s call sits inside it for ~11 consecutive\n    # 0.5 s-strided positions, and if all of them clear CAND_FLOOR, stitch_time merges them\n    # (0.5 < the default gap of 0.505 s) into one event spanning 0.5(R-1)+1 = 6 s. Truth\n    # spans are 1 s, so IoU = 1/6 = 0.17, under IOU_THR=0.30 -- no match. Every merged-run\n    # length R in [6,19] is unmatchable this way, and the default MAX_EVENT_SEC=10 only\n    # rescues R>=20 (truncated to the peak window). This is a geometric hazard, not a\n    # proven outcome -- whether real above-floor runs actually land in [6,19] depends on\n    # the empirical density of the score field, which was never measured before this fix\n    # landed (found by audit 28 Aug 2026, before Cell 5 ever ran). If it did land there\n    # for most calls, the failure mode is silent and severe: near-zero true positives push\n    # thr toward 1.0, drive FA/h toward 0 and AUC toward the 0.500 all-ties value, which\n    # verdict() would report as \"Perch alone is WORSE -- stage-1's localisation is doing\n    # real work\": our own pre-registered prediction, confirmed by an artefact rather than\n    # a result. Capping at SEG_SEC removes the hazard rather than characterising it: it\n    # collapses every run to a 1 s span at its peak window (R=1 is already 1 s and passes\n    # through untouched), which is also the honest semantic -- a 5 s window cannot localise\n    # better than its argmax. gap_s is passed explicitly so widening stride_s past 0.505 s\n    # cannot silently flip merging off. Found by audit 28 Aug 2026, before this ever ran.\n    ev = stitch_time(np.asarray(times), v, gap_s=max(stride_s * 1.5, 1e-6),\n                     max_event_sec=SEG_SEC)\n    return [(s, e, vi, vi) for s, e, vi in ev]\n\nFUSIONS = {\n    \"stage-1 only  (p)\":  lambda p, v: p,\n    \"perch only    (v)\":  lambda p, v: v,\n    \"p * v\":              lambda p, v: p * v,\n    \"p * sqrt(v)\":        lambda p, v: p * np.sqrt(v),\n    \"sqrt(p) * v\":        lambda p, v: np.sqrt(p) * v,\n    \"geo mean sqrt(p*v)\": lambda p, v: np.sqrt(p * v),\n}\n\n# ------------------------------------------------------------------- metrics\ndef match(events, truth, iou_thresh):\n    pairs = sorted(((iou(e[:2], t), i, j) for i, e in enumerate(events)\n                    for j, t in enumerate(truth)), reverse=True)\n    used_e, used_t = set(), set()\n    for score, i, j in pairs:\n        if score <= iou_thresh: break\n        if i in used_e or j in used_t: continue\n        used_e.add(i); used_t.add(j)\n    return [(e[2], i in used_e) for i, e in enumerate(events)]\n\ndef match_truth(events, truth, iou_thresh):\n    pairs = sorted(((iou(e[:2], t), i, j) for i, e in enumerate(events)\n                    for j, t in enumerate(truth)), reverse=True)\n    ue, ut, out = set(), set(), [0.0] * len(truth)\n    for score, i, j in pairs:\n        if score <= iou_thresh: break\n        if i in ue or j in ut: continue\n        ue.add(i); ut.add(j); out[j] = events[i][2]\n    return out\n\ndef fp_per_tp(scored, n_truth, target_recall=0.5):\n    arr = sorted(scored, key=lambda x: -x[0]); tp = fp = 0\n    for score, hit in arr:\n        if hit: tp += 1\n        else:   fp += 1\n        if tp / n_truth >= target_recall:\n            return fp / tp if tp else float('inf')\n    return float('inf')\n\ndef summarise(scored, n_truth):\n    if not scored or not n_truth:\n        return dict(ap=0.0, f1=0.0, p=0.0, r=0.0, thr=1.0, thr_r90=1.0, n=len(scored))\n    arr = sorted(scored, key=lambda x: -x[0])\n    tp = fp = 0; ap = 0.0; prev_r = 0.0; best = (0.0, 0.0, 0.0, 1.0); thr_r90 = None\n    for score, hit in arr:\n        if hit: tp += 1\n        else:   fp += 1\n        pr = tp / (tp + fp); rc = tp / n_truth\n        ap += pr * (rc - prev_r); prev_r = rc\n        f1 = 2*pr*rc/(pr+rc) if pr + rc else 0.0\n        if f1 > best[0]: best = (f1, pr, rc, score)\n        # Recall-first operating point (N4): the HIGHEST threshold still achieving >=90%\n        # recall, i.e. best precision subject to a recall floor. Scores descend, so the\n        # first crossing is that threshold. Setting it to the score just BEFORE the\n        # crossing (the earlier bug) yields a point that never actually reaches 90%.\n        if thr_r90 is None and rc >= 0.90: thr_r90 = score\n    if thr_r90 is None: thr_r90 = arr[-1][0]   # 90% unreachable -> loosest available\n    return dict(ap=ap, f1=best[0], p=best[1], r=best[2], thr=best[3],\n                thr_r90=thr_r90, n=len(scored))\n\ndef _auc(pos, neg):\n    \"\"\"P(random target scores above random decoy), ties 0.5. Threshold-FREE.\n    Same metric family BirdCLEF 2024 adopted as its official score (macro ROC-AUC):\n    0.50 = cannot tell target from decoy, 1.00 = perfect species discrimination.\"\"\"\n    pos, neg = np.asarray(pos, float), np.asarray(neg, float)\n    if not len(pos) or not len(neg): return float('nan')\n    gt = float((pos[:, None] > neg[None, :]).sum())\n    eq = float((pos[:, None] == neg[None, :]).sum())\n    return (gt + 0.5 * eq) / (len(pos) * len(neg))\n\n# ------------------------------------------------------------------- passes\n# detect_fn is swappable (default detect_pv) so CELL 5 can run the identical planting /\n# matching logic against a different detector -- e.g. detect_perch_only -- and get a\n# directly comparable number rather than a separately-implemented, possibly-diverging one.\ndef eval_pass(enc, ctx, seed, detect_fn=detect_pv):\n    r = np.random.default_rng(seed); cache = []\n    for bed in eval_beds:\n        y = bed.copy(); L = len(y); spans, planted, snrs = [], [], []\n        for k in range(CALLS_PER_BED):\n            snr = SNR_LEVELS[k % len(SNR_LEVELS)]\n            placed = None\n            for _ in range(80):\n                pos = int(r.integers(0, L-CLEN))\n                if any(pos < e+int(0.4*SR) and pos+CLEN > s-int(0.4*SR) for s, e in spans): continue\n                placed = pos; break\n            if placed is None: continue\n            c = ctx[\"target_pool\"][int(r.integers(len(ctx[\"target_pool\"])))][:CLEN]\n            if len(c) < CLEN: c = np.pad(c, (0, CLEN-len(c)))\n            g = rms(y[placed:placed+CLEN])*(10**(snr/20))/(rms(c)+1e-9)\n            y[placed:placed+CLEN] += c.astype(\"float32\")*g\n            spans.append((placed, placed+CLEN))\n            planted.append((placed/SR, (placed+CLEN)/SR)); snrs.append(snr)\n        cache.append((detect_fn(enc, ctx, y, r), planted, snrs))\n    return cache\n\ndef spec_pass(enc, ctx, seed, snr_db=6.0, detect_fn=detect_pv):\n    \"\"\"Specificity control. snr_db is now a parameter: v1 fixed it at +6 dB, so species\n    discrimination had only ever been measured on loud, easy calls (roadmap 3.2).\"\"\"\n    r = np.random.default_rng(seed); cache = []\n    for bed in eval_beds:\n        y = bed.copy(); L = len(y); spans, items = [], []\n        plan = [(\"target\", ctx[\"target_pool\"])]*6\n        for n_, p_ in ctx[\"decoy_pools\"].items(): plan += [(n_, p_)]*3\n        for kind, pool in plan:\n            placed = None\n            for _ in range(80):\n                pos = int(r.integers(0, L-CLEN))\n                if any(pos < e+int(0.4*SR) and pos+CLEN > s-int(0.4*SR) for s, e in spans): continue\n                placed = pos; break\n            if placed is None or not pool: continue\n            c = pool[int(r.integers(len(pool)))][:CLEN]\n            if len(c) < CLEN: c = np.pad(c, (0, CLEN-len(c)))\n            g = rms(y[placed:placed+CLEN])*(10**(snr_db/20))/(rms(c)+1e-9)\n            y[placed:placed+CLEN] += c.astype(\"float32\")*g\n            spans.append((placed, placed+CLEN))\n            items.append((kind, (placed/SR, (placed+CLEN)/SR)))\n        cache.append((detect_fn(enc, ctx, y, r), items))\n    return cache\n\ndef neg_pass(enc, ctx, seed, detect_fn=detect_pv):\n    \"\"\"PURE NEGATIVE control -- unmodified beds, zero target calls planted.\n    Every detection here is a false alarm by construction. This is the deployment-\n    relevant number (false alarms per hour of clean audio) and v1 never measured it.\"\"\"\n    r = np.random.default_rng(seed + 500); cache = []\n    for bed in eval_beds:\n        cache.append(detect_fn(enc, ctx, bed.copy(), r))\n    return cache\n\n# ------------------------------------------------------------------- scoring\ndef score_eval(fuse, cache):\n    scored, n_truth, by_snr = [], 0, []\n    for dets, planted, snrs in cache:\n        ev = [(s, e, float(fuse(p, v))) for s, e, p, v in dets]\n        scored += match(ev, planted, IOU_THR); n_truth += len(planted)\n        by_snr += list(zip(snrs, match_truth(ev, planted, IOU_THR)))\n    m = summarise(scored, n_truth)\n    m[\"fp_per_tp@R50\"] = fp_per_tp(scored, n_truth, 0.5)\n    m[\"recall_by_snr\"] = {s: float(np.mean([sc >= m[\"thr\"] for ss, sc in by_snr if ss == s]))\n                          for s in SNR_LEVELS}\n    return m\n\ndef score_spec(fuse, thr, cache, decoy_names):\n    tg, dc = [], {k: [] for k in decoy_names}\n    for dets, items in cache:\n        ev = [(s, e, float(fuse(p, v))) for s, e, p, v in dets]\n        for kind, span in items:\n            best = max((iou(e[:2], span), e[2]) for e in ev) if ev else (0.0, 0.0)\n            sc = best[1] if best[0] > IOU_THR else 0.0\n            (tg if kind == \"target\" else dc[kind]).append(sc)\n    t = float(np.mean([s >= thr for s in tg])) if tg else 0.0\n    d = {k: float(np.mean([s >= thr for s in v])) for k, v in dc.items() if v}\n    md = float(np.mean(list(d.values()))) if d else 0.0\n    all_dc = [s for v in dc.values() for s in v]\n    return t, md, (md/t if t else float('nan')), d, _auc(tg, all_dc)\n\ndef spec_coverage(cache):\n    \"\"\"-> fraction of planted items stage 1 LOCALISED at all (IoU > IOU_THR), by kind.\n\n    Only the SNR sweep (roadmap 3.2) needs this, and there it is not optional. score_spec\n    assigns 0.0 to any planted item that was never localised, and _auc counts a 0.0-vs-0.0\n    pair as a tie worth 0.5. So an AUC drifting toward 0.50 at low SNR is ambiguous between\n    two different findings: \"target and decoy became indistinguishable\" and \"nothing was\n    detected, so everything tied\". Coverage separates them. Reporting the SNR curve without\n    it would invite exactly the misreading the curve is meant to settle.\n\n    Fusion-independent by construction -- a fusion rule changes an event's score, never its\n    span -- so this is computed once per cache rather than once per fusion rule.\n    \"\"\"\n    n = {\"target\": 0, \"decoy\": 0}; hit = {\"target\": 0, \"decoy\": 0}\n    for dets, items in cache:\n        for kind, span in items:\n            k = \"target\" if kind == \"target\" else \"decoy\"\n            n[k] += 1\n            best = max((iou((s, e), span) for s, e, _p, _v in dets), default=0.0)\n            hit[k] += int(best > IOU_THR)\n    return {k: (hit[k]/n[k] if n[k] else float('nan')) for k in n}\n\ndef score_spec_detected(fuse, cache):\n    \"\"\"-> (AUC over LOCALISED items only, n_target_localised, n_decoy_localised).\n\n    The companion to spec_coverage, and the direct measurement it was a proxy for.\n    score_spec's AUC mixes two effects that pull in the same direction at low SNR: genuine\n    species confusion, and stage 1 failing to localise the call at all -- which forces a\n    0.0 and drags the AUC toward 0.50 for a reason that has nothing to do with species\n    identity. Excluding unlocalised items (rather than scoring them 0.0) separates the two,\n    so \"specificity degraded\" and \"detection degraded\" can be told apart from the numbers\n    instead of argued about afterwards.\n\n    Returns nan when either side is empty -- at low enough SNR that is the honest answer,\n    and it must not be silently averaged in as if it were a measurement.\n    \"\"\"\n    tg, dc = [], []\n    for dets, items in cache:\n        ev = [(s, e, float(fuse(p, v))) for s, e, p, v in dets]\n        for kind, span in items:\n            best = max((iou(e[:2], span), e[2]) for e in ev) if ev else (0.0, 0.0)\n            if best[0] <= IOU_THR: continue      # never localised -> excluded, NOT zeroed\n            (tg if kind == \"target\" else dc).append(best[1])\n    return _auc(tg, dc), len(tg), len(dc)\n\ndef score_neg(fuse, thr, cache):\n    \"\"\"-> false alarms per hour on audio containing no target calls.\"\"\"\n    fa = sum(sum(1 for s, e, p, v in dets if float(fuse(p, v)) >= thr) for dets in cache)\n    hours = (len(cache) * BED_SEC) / 3600.0\n    return fa / hours if hours else float('nan')\n\n# ------------------------------- stage-1 encoder (ALL targets excluded -- the bug fix)\ndef build_banks(exclude, features=FEATURES, seed=0, max_species=40):\n    \"\"\"One bank per front-end, sharing a single pass over the audio.\n\n    An encoder trained on PCEN cannot be evaluated on log-Mel inputs, so the front-end\n    cannot be a same-encoder ladder rung -- it needs its own encoder. But the expensive\n    part is decoding audio, not featurising it, so each recording is loaded ONCE and\n    featurised for every front-end. Cost is one extra encoder training, not a second\n    pass over the dataset.\n    \"\"\"\n    rng = np.random.default_rng(seed)\n    sps = [s for s in meta.primary_label.unique() if s not in exclude]\n    rng.shuffle(sps)\n    banks = {f: {} for f in features}\n    for sp in sps:\n        if len(banks[features[0]]) >= max_species: break\n        segs = {f: [] for f in features}\n        for fn in meta[meta.primary_label == sp][\"filename\"].tolist()[:12]:\n            try:\n                y = ld(os.path.join(AUD, fn), 15)\n                c = loud_off(y)//HOP + SEG_FR//2\n                for f in features:\n                    segs[f].append(crop(feat(y, f), c))\n            except Exception:\n                pass\n        if len(segs[features[0]]) >= 8:\n            for f in features:\n                banks[f][sp] = segs[f]\n    return banks\n\nprint(f\"\\nbuilding encoder banks for {list(FEATURES)}, excluding all {len(TARGET_SET)} \"\n      f\"targets: {sorted(TARGET_SET)}\")\nbanks = build_banks(TARGET_SET, FEATURES, seed=0)\nfor _f, _b in banks.items():\n    assert not (TARGET_SET & set(_b)), \\\n        f\"LEAKAGE: target species present in {_f} encoder bank: {TARGET_SET & set(_b)}\"\n    print(f\"  {_f:<7} bank: {len(_b)} species, none of them targets (assertion passed)\")\n\n# Keyed by (feature, seed): a rung that switches the front-end must switch the encoder\n# with it, or it is measuring an encoder/feature mismatch rather than the front-end.\nencoders = {}\nfor _f in FEATURES:\n    for es in ENCODER_SEEDS:\n        print(f\"training stage-1 encoder (feature={_f}, seed={es})...\")\n        torch.manual_seed(es)\n        encoders[(_f, es)] = train_enc(banks[_f], np.random.default_rng(es))\n\nif SKIP_MAIN_LOOP:\n    print(\"SKIP_MAIN_LOOP=True -- stopping after CELL 2's encoder training.\")\n    print(\"res/wide/the per-species report do not exist this run. Proceed to CELL 4;\")\n    print(\"its own checkpointing resumes the ladder at the first unfinished config.\")\nelse:\n    # ------------------------------------------------------------------- main loop\n    rows, t_start = [], time.time()\n    for sp in TARGET_LIST:\n        print(f\"\\n{'='*78}\\nTARGET: {sp}\\n{'='*78}\")\n        t_sp = time.time()\n        try:\n            ctx = build_target(sp)\n        except Exception as ex:\n            print(f\"  SKIP {sp}: {type(ex).__name__}: {ex}\"); continue\n\n        for es in ENCODER_SEEDS:\n            enc = encoders[(FEATURE, es)]\n            per_seed = {name: [] for name in FUSIONS}\n            for sd in SEEDS:\n                t0 = time.time()\n                ec = eval_pass(enc, ctx, sd)\n                sc = spec_pass(enc, ctx, sd)\n                nc = neg_pass(enc, ctx, sd)\n                for name, fuse in FUSIONS.items():\n                    m = score_eval(fuse, ec)\n                    t, md, ratio, per, auc = score_spec(fuse, m[\"thr\"], sc, ctx[\"DECOYS\"])\n                    per_seed[name].append(dict(\n                        ap=m[\"ap\"], f1=m[\"f1\"], p=m[\"p\"], r=m[\"r\"],\n                        fptp=m[\"fp_per_tp@R50\"], auc=auc, ratio=ratio,\n                        fa_per_hr=score_neg(fuse, m[\"thr\"], nc),\n                        fa_per_hr_r90=score_neg(fuse, m[\"thr_r90\"], nc)))\n                print(f\"  seed {sd} done in {time.time()-t0:.0f}s\")\n\n            for name in FUSIONS:\n                agg = {k: float(np.mean([d[k] for d in per_seed[name]])) for k in per_seed[name][0]}\n                sds = {k + \"_sd\": float(np.std([d[k] for d in per_seed[name]], ddof=1))\n                       if len(SEEDS) > 1 else 0.0 for k in per_seed[name][0]}\n                rows.append({**ctx[\"covariates\"], \"encoder_seed\": es, \"fusion\": name,\n                             \"feature\": FEATURE, \"n_seeds\": len(SEEDS),\n                             \"use_timefilter\": USE_TIMEFILTER, \"use_hard_neg\": USE_HARD_NEG,\n                             **agg, **sds})\n        print(f\"  [{sp}] total {time.time()-t_sp:.0f}s\")\n\n    res = pd.DataFrame(rows)\n    res.to_csv(\"/kaggle/working/argus_multispecies_results.csv\", index=False)\n    print(f\"\\nwrote {len(res)} rows -> argus_multispecies_results.csv \"\n          f\"| wall clock {(time.time()-t_start)/60:.1f} min\")\n\n    # ------------------------------------------------------------------- headline table\n    BASE, CAND = \"stage-1 only  (p)\", \"perch only    (v)\"\n    print(f\"\\n{'='*100}\\nSPECIFICITY AUC BY SPECIES  (threshold-free; 0.50 = target \"\n          f\"indistinguishable from decoys)\\n{'='*100}\")\n    print(f\"{'species':<12}{'n_rec':>6}{'decoy_sim':>10}{'stereo':>8}\"\n          f\"{'AUC stage-1':>13}{'AUC verified':>14}{'delta':>8}{'F1 s1':>7}{'F1 ver':>8}\")\n    print(\"-\"*100)\n    # Average over encoder seeds first, so a multi-encoder-seed run reports the mean rather\n    # than silently taking whichever row happened to land first.\n    NUM = [\"auc\", \"f1\", \"fa_per_hr\", \"fa_per_hr_r90\", \"n_recordings\",\n           \"decoy_sim_top1\", \"stereotypy\"]\n    piv = (res[res.fusion.isin([BASE, CAND])]\n           .groupby([\"species\", \"fusion\"], sort=False)[NUM].mean().reset_index())\n    wide = piv.pivot(index=\"species\", columns=\"fusion\")\n\n    for sp in [s for s in TARGET_LIST if s in wide.index]:\n        r = wide.loc[sp]\n        print(f\"{sp:<12}{int(r[('n_recordings', BASE)]):>6}{r[('decoy_sim_top1', BASE)]:>10.3f}\"\n              f\"{r[('stereotypy', BASE)]:>8.3f}{r[('auc', BASE)]:>13.3f}{r[('auc', CAND)]:>14.3f}\"\n              f\"{r[('auc', CAND)] - r[('auc', BASE)]:>+8.3f}\"\n              f\"{r[('f1', BASE)]:>7.3f}{r[('f1', CAND)]:>8.3f}\")\n\n    # ---------------------------------------------------------- encoder-seed spread\n    # The table above averages over encoder seeds -- which hides exactly the variance\n    # ENCODER_SEEDS was widened to measure. Same bank, same data, same eval seeds; the only\n    # thing differing is the encoder training run. Anything smaller than this spread is not\n    # a real effect, no matter what a within-run SE says about it.\n    if len(ENCODER_SEEDS) > 1:\n        print(f\"\\n{'='*100}\\nENCODER-SEED SPREAD -- pure training-run variance \"\n              f\"(same bank, same data, {len(ENCODER_SEEDS)} encoder seeds)\\n{'='*100}\")\n        es_piv = (res[res.fusion.isin([BASE, CAND])]\n                  .groupby([\"species\", \"fusion\", \"encoder_seed\"])[\"auc\"].mean().reset_index())\n        worst = 0.0\n        for sp in [s for s in TARGET_LIST if s in set(es_piv.species)]:\n            for fus, tag in ((BASE, \"stage-1\"), (CAND, \"verified\")):\n                g = es_piv[(es_piv.species == sp) & (es_piv.fusion == fus)].sort_values(\"encoder_seed\")\n                if g.empty: continue\n                v = g.auc.to_numpy(); spread = float(v.max() - v.min())\n                worst = max(worst, spread)\n                print(f\"  {sp:<10} {tag:<9} \" + \"  \".join(f\"{x:.3f}\" for x in v)\n                      + f\"   | spread {spread:+.3f}  sd {v.std(ddof=1):.3f}\")\n        print(f\"\\n  Largest spread from encoder training alone: {worst:.3f}\")\n        print(f\"  Treat any AUC difference below ~{worst:.3f} as noise, not signal.\")\n\n    if len(wide) > 1:\n        d = (wide[(\"auc\", CAND)] - wide[(\"auc\", BASE)]).to_numpy()\n        se = d.std(ddof=1)/np.sqrt(len(d))\n        print(f\"\\nACROSS SPECIES (n={len(d)}): mean AUC gain {d.mean():+.3f}, SE {se:.3f}, \"\n              f\"delta/SE {d.mean()/se if se else float('nan'):.2f}   (|d/SE| < 2 is NOT evidence)\")\n        print(f\"  verification helped:            {int((d > 0).sum())}/{len(d)} species\")\n        print(f\"  below chance WITHOUT verifying: \"\n              f\"{int((wide[('auc', BASE)] < 0.5).sum())}/{len(d)} species\")\n        print(f\"  below chance WITH verifying:    \"\n              f\"{int((wide[('auc', CAND)] < 0.5).sum())}/{len(d)} species\")\n        # The pre-registered outcome test (roadmap sec.2): does the original single-species\n        # finding generalise, hold only conditionally, or fail? Reported either way.\n        frac = (d > 0).mean()\n        print(\"  -> outcome \" + (\"A (generalises)\" if frac >= 0.8 else\n                                 \"B (conditional -- find what separates them)\" if frac >= 0.4 else\n                                 \"C/D (does NOT generalise -- rebuild narrative around the boundary)\"))\n\n    print(f\"\\nFALSE ALARMS PER HOUR on clean audio (no target present)\")\n    print(f\"{'species':<12}{'best-F1 thr':>24}{'recall-first thr':>26}\")\n    for sp in [s for s in TARGET_LIST if s in wide.index]:\n        r = wide.loc[sp]\n        print(f\"  {sp:<10}{r[('fa_per_hr', BASE)]:>10.1f} ->{r[('fa_per_hr', CAND)]:>8.1f}\"\n              f\"{r[('fa_per_hr_r90', BASE)]:>16.1f} ->{r[('fa_per_hr_r90', CAND)]:>8.1f}\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"2aa7c805","cell_type":"markdown","source":"## Cell 3 — Shot-Selection & Shot-Count Sweeps\n\n`arguswala_sweeps.py`","metadata":{}},{"id":"87c562ee","cell_type":"code","source":"# ===== CELL 3 / SWEEPS -- shot selection, shot count, front-end =====\n# Run AFTER arguswala_v2.py (reuses its trained `encoders`, scoring fns and eval passes).\n#\n# Five experiments the literature says matter and nobody has published answers for on\n# this data:\n#\n#  N1  SHOT-SELECTION STRATEGY.  Nolasco et al. (2023) sec.4.1: unrepresentative support\n#      sets were the single failure cause their expert annotators raised most often --\n#      \"the five shots ... are not representative of the range of possible [calls]\".\n#      DCASE always takes the FIRST five. ARGUS v1 always took the highest-RATED five.\n#      Nobody has measured what that choice is worth. Four strategies, same everything else.\n#\n#  S2  SHOT COUNT.  ARGUS is a few-shot project that has never varied the number of shots.\n#      First thing a judge asks. Where the curve plateaus is a deployable field number.\n#\n#  S3  SPECIFICITY vs SNR.  Every specificity number this project has reported was\n#      measured at a fixed +6 dB -- loud, close calls. Roadmap 3.2 calls this a real hole,\n#      with a pre-registered prediction attached. A deployment cannot choose its SNR.\n#\n#  S4  DECOY DIFFICULTY.  Decoys have always been the top-N_DECOYS hardest confusers, on\n#      purpose -- never measured against the alternative. Roadmap 3.3.\n#\n#  N2  FRONT-END.  Liang et al. (DCASE 2024 baseline) Table III: PCEN has the best\n#      PRECISION of any front-end tested (68.0%), log-Mel the best F1 (63.7 vs 60.0).\n#      ARGUS uses PCEN. Confirm or contradict that trade on Western Ghats data.\n#      NOTE: changing FEATURE changes the encoder bank too, so this needs a fresh\n#      encoder -- it is NOT a free re-score. Budget one extra encoder training.\nimport numpy as np, pandas as pd, time, torch\n\n# 29 Aug 2026: widened from 3 species x 2 seeds (= 6 measurements per condition) to\n# 5 x 3 (= 15). The old figure sat below the 10-per-condition bar the IRIS/ISEF rubric\n# names explicitly as a pitfall (\"sample size of three ... won't survive statistical\n# testing\"), and the sweeps were the weakest link in the project on that criterion -- the\n# centrepiece already carries n=10 species for its paired test, the confounds carry 5.\n#\n# Species were widened in preference to seeds at equal cost: 5x3 and 3x5 both cost 2.5x\n# the old run, but extra species test whether a sweep's verdict GENERALISES, where extra\n# seeds only tighten the noise estimate on the same three. It also aligns SWEEP_SPECIES\n# with CONFOUND_SPECIES (both TARGET_LIST[:5]), so Cell 3 and Cell 5 now describe the same\n# species set rather than two overlapping ones.\n#\n# Projected cost from run 3's timings (N1 3350s + S2 3331s + S3 2869s + S4 2614s = 3.38h\n# at 3x2): ~8.5h at 5x3, plus ~10min setup/encoders. Fits the 12h ceiling with ~3h spare.\n# Each sweep writes its CSV on completion, so a session killed during S4 still keeps\n# N1/S2/S3.\nSWEEP_SPECIES = TARGET_LIST[:5]      # sweeps are O(strategies x species x seeds); keep narrow\nSWEEP_SEEDS   = (7, 8, 9)            # fewer than the main run -- these are relative comparisons\nBASE, CAND    = \"stage-1 only  (p)\", \"perch only    (v)\"\nENC           = encoders[(FEATURE, ENCODER_SEEDS[0])]   # encoders are keyed by front-end\n\n# 28 Aug 2026, after run 4 (12h ceiling, killed mid-ladder, Cell 5 never reached):\n# all four sweeps below (N1/S2/S3/S4) have now independently replicated twice -- 16 Aug\n# and 25 Aug -- with the same qualitative verdict both times (audited against run 4's own\n# console log the same day; see the lab notebook Day 34). SKIP_SWEEPS=True skips this\n# cell's ~3.4h of already-answered compute so a Cell-5-focused session goes straight to\n# the one thing genuinely unrun. Same revert-after-use discipline as SKIP_MAIN_LOOP /\n# SKIP_LADDER: leave False for a normal run, or when re-verifying a sweep is the point.\n#\n# 29 Aug 2026 -- THREE SESSIONS, RUN SEPARATELY. Never combine this cell with another.\n#   [DONE 29 Aug] P2 re-run: SKIP_SWEEPS=True, RUN_P2=True -> ran 3h40m, settled P2\n#                 (stage-1 FLAT, verified FALLS -0.050, negative in 5/5 species).\n#   [THIS BUILD]  widened sweeps: SKIP_SWEEPS=False, everything else off -> ~8.7h.\n#   [PENDING]     distillation: Cell 6 SKIP_DISTILL=False, this cell back to True -> ~3h.\n# The sweeps alone project to ~8.7h against a 12h ceiling; adding either other session\n# overruns it, which is how an earlier run was lost mid-ladder.\nSKIP_SWEEPS = True\n\ndef run_one(sp, n_shot, strategy, seeds=SWEEP_SEEDS, decoy_mode=\"hard\"):\n    \"\"\"One (species, n_shot, strategy, decoy_mode) cell -> mean metrics over seeds.\n    decoy_mode defaults to 'hard' (roadmap 3.3's meaning of \"unchanged\") so N1/S2 above\n    are byte-for-byte the same run they always were; only S4 below passes anything else.\"\"\"\n    try:\n        ctx = build_target(sp, n_shot=n_shot, strategy=strategy, verbose=False,\n                            decoy_mode=decoy_mode)\n    except Exception as ex:\n        print(f\"    skip {sp} n_shot={n_shot} {strategy} {decoy_mode}: \"\n              f\"{type(ex).__name__}: {ex}\")\n        return None\n    acc = {BASE: [], CAND: []}\n    for sd in seeds:\n        ec = eval_pass(ENC, ctx, sd)\n        sc = spec_pass(ENC, ctx, sd)\n        for name in (BASE, CAND):\n            m = score_eval(FUSIONS[name], ec)\n            _, _, _, _, auc = score_spec(FUSIONS[name], m[\"thr\"], sc, ctx[\"DECOYS\"])\n            acc[name].append((auc, m[\"f1\"], m[\"p\"], m[\"r\"]))\n    out = dict(species=sp, n_shot=n_shot, strategy=strategy, decoy_mode=decoy_mode,\n               n_heldout=ctx[\"covariates\"][\"n_heldout_calls\"],\n               decoy_sim_top1=ctx[\"covariates\"][\"decoy_sim_top1\"])\n    for name, tag in ((BASE, \"s1\"), (CAND, \"ver\")):\n        a = np.array(acc[name], dtype=float)\n        out[f\"auc_{tag}\"] = a[:, 0].mean(); out[f\"auc_{tag}_sd\"] = a[:, 0].std(ddof=1) if len(a) > 1 else 0.0\n        out[f\"f1_{tag}\"]  = a[:, 1].mean()\n        out[f\"prec_{tag}\"] = a[:, 2].mean(); out[f\"rec_{tag}\"] = a[:, 3].mean()\n    return out\n\nif not SKIP_SWEEPS:\n    # ------------------------------------------------------------------ N1: shot selection\n    print(f\"\\n{'='*84}\\nN1  SHOT-SELECTION STRATEGY  (n_shot={N_SHOT} held constant)\\n{'='*84}\")\n    t0 = time.time(); rows_n1 = []\n    for strategy in (\"first\", \"rated\", \"random\", \"diverse\"):\n        for sp in SWEEP_SPECIES:\n            r = run_one(sp, N_SHOT, strategy)\n            if r: rows_n1.append(r); print(f\"  {strategy:<8} {sp:<12} \"\n                                           f\"AUC {r['auc_s1']:.3f} -> {r['auc_ver']:.3f}   \"\n                                           f\"F1 {r['f1_s1']:.3f} -> {r['f1_ver']:.3f}\")\n    n1 = pd.DataFrame(rows_n1)\n    if len(n1):\n        n1.to_csv(\"/kaggle/working/argus_sweep_shot_strategy.csv\", index=False)\n        print(f\"\\n{'strategy':<10}{'AUC stage-1':>13}{'AUC verified':>14}{'F1 stage-1':>12}{'F1 verified':>13}\")\n        print(\"-\"*62)\n        for s, g in n1.groupby(\"strategy\", sort=False):\n            print(f\"{s:<10}{g.auc_s1.mean():>13.3f}{g.auc_ver.mean():>14.3f}\"\n                  f\"{g.f1_s1.mean():>12.3f}{g.f1_ver.mean():>13.3f}\")\n        best = n1.groupby(\"strategy\").auc_ver.mean().idxmax()\n        spread = n1.groupby(\"strategy\").auc_ver.mean()\n        print(f\"\\nbest strategy by verified AUC: {best}  \"\n              f\"(spread across strategies: {spread.max()-spread.min():.3f} AUC)\")\n        print(\"If that spread exceeds the verification gain itself, then HOW the five shots are\"\n              \"\\nchosen matters more than the architecture -- which is a publishable observation\"\n              \"\\nand exactly the open problem Nolasco et al. sec.4.1 raise.\")\n    print(f\"[N1 took {(time.time()-t0)/60:.1f} min]\")\n\n    # ------------------------------------------------------------------ S2: shot count\n    print(f\"\\n{'='*84}\\nS2  SHOT COUNT  (strategy='{SHOT_STRATEGY}' held constant)\\n{'='*84}\")\n    t0 = time.time(); rows_s2 = []\n    for n_shot in (1, 3, 5, 10):\n        for sp in SWEEP_SPECIES:\n            r = run_one(sp, n_shot, SHOT_STRATEGY)\n            if r: rows_s2.append(r); print(f\"  n_shot={n_shot:<3} {sp:<12} \"\n                                           f\"AUC {r['auc_s1']:.3f} -> {r['auc_ver']:.3f}   \"\n                                           f\"F1 {r['f1_s1']:.3f} -> {r['f1_ver']:.3f}\")\n    s2 = pd.DataFrame(rows_s2)\n    if len(s2):\n        s2.to_csv(\"/kaggle/working/argus_sweep_shot_count.csv\", index=False)\n        print(f\"\\n{'n_shot':<8}{'AUC stage-1':>13}{'AUC verified':>14}{'F1 verified':>13}\")\n        print(\"-\"*48)\n        for n, g in s2.groupby(\"n_shot\"):\n            print(f\"{n:<8}{g.auc_s1.mean():>13.3f}{g.auc_ver.mean():>14.3f}{g.f1_ver.mean():>13.3f}\")\n        print(\"\\nRead the plateau, not the peak: the smallest n_shot within noise of the best is\"\n              \"\\nthe number a field team actually has to collect.\")\n    print(f\"[S2 took {(time.time()-t0)/60:.1f} min]\")\n\n    # ------------------------------------------------------------- S3: specificity vs SNR\n    # Roadmap 3.2 -- the last unmeasured axis in the scenario matrix. spec_pass() has always\n    # planted at a fixed +6 dB, so EVERY species-discrimination number this project has ever\n    # reported -- including the 0.751 centrepiece -- describes loud, close calls only. A field\n    # deployment does not get to choose the SNR.\n    #\n    # PRE-REGISTERED PREDICTION (roadmap 3.2, written before this ever ran):\n    #   \"the verifier's advantage shrinks as calls get fainter.\"\n    #   If true  -> an honest limitation with a number attached, and the centrepiece result\n    #               must be quoted as a HIGH-SNR result from then on.\n    #   If false -> a stronger claim than the project currently makes.\n    # Both outcomes get reported. Nothing about the verdict is decided after seeing the curve.\n    SNR_SWEEP_SPECIES = SWEEP_SPECIES     # same 3 as N1/S2, so all three sweeps stay comparable\n    SNR_SWEEP_SEEDS   = SWEEP_SEEDS\n    SNR_SWEEP_LEVELS  = SNR_LEVELS        # (9, 6, 3, 0, -3, -6); trim first if budget is tight\n    NOISE_FLOOR       = 0.025             # measured in-session Day 26-27. Nothing below it is real.\n\n    def run_snr(sp, seeds=SNR_SWEEP_SEEDS, levels=SNR_SWEEP_LEVELS):\n        \"\"\"One species -> one row per (seed, SNR level, fusion).\n\n        eval_pass runs ONCE per seed rather than once per level. It exists here only to set the\n        operating threshold, and its own planting protocol is SNR-independent by design (it\n        rotates through SNR_LEVELS internally). Re-running it per level would cost 6x for an\n        identical threshold -- and would also be the wrong experiment. Calibrating once and then\n        varying the call level IS the deployment case: you fix a detector's operating point, and\n        the forest hands you whatever it hands you.\n        \"\"\"\n        try:\n            ctx = build_target(sp, verbose=False)\n        except Exception as ex:\n            print(f\"    skip {sp}: {type(ex).__name__}: {ex}\")\n            return []\n        rows = []\n        for sd in seeds:\n            ec  = eval_pass(ENC, ctx, sd)\n            thr = {name: score_eval(FUSIONS[name], ec)[\"thr\"] for name in (BASE, CAND)}\n            for snr in levels:\n                sc  = spec_pass(ENC, ctx, sd, snr_db=float(snr))\n                cov = spec_coverage(sc)          # fusion-independent, so computed once per level\n                for name, tag in ((BASE, \"s1\"), (CAND, \"ver\")):\n                    t, md, ratio, _per, auc = score_spec(FUSIONS[name], thr[name], sc, ctx[\"DECOYS\"])\n                    auc_det, n_td, n_dd = score_spec_detected(FUSIONS[name], sc)\n                    rows.append(dict(species=sp, seed=sd, snr_db=int(snr), fusion=tag,\n                                     auc=auc, auc_det=auc_det, n_tgt_det=n_td, n_dec_det=n_dd,\n                                     tgt_rate=t, decoy_rate=md, ratio=ratio,\n                                     cov_target=cov[\"target\"], cov_decoy=cov[\"decoy\"],\n                                     thr=float(thr[name])))\n            print(f\"    {sp:<12} seed {sd} done ({len(levels)} levels)\")\n        return rows\n\n    _np = len(SNR_SWEEP_SPECIES) * len(SNR_SWEEP_SEEDS) * (1 + len(SNR_SWEEP_LEVELS))\n    print(f\"\\n{'='*84}\\nS3  SPECIFICITY vs SNR  (roadmap 3.2)\\n{'='*84}\")\n    print(f\"{len(SNR_SWEEP_SPECIES)} species x {len(SNR_SWEEP_SEEDS)} seeds x \"\n          f\"({len(SNR_SWEEP_LEVELS)} levels + 1 threshold pass) = {_np} detection passes \"\n          f\"~= {_np*1.1:.0f} min at the measured 1.1 min/pass.\")\n    print(\"PRE-REGISTERED: the verification gain is predicted to SHRINK as SNR falls.\")\n    print(\"Reported either way -- a flat curve is the stronger result, not a failed run.\\n\")\n\n    t0 = time.time(); rows_s3 = []\n    for sp in SNR_SWEEP_SPECIES:\n        rows_s3 += run_snr(sp)\n    s3 = pd.DataFrame(rows_s3)\n\n    if len(s3):\n        s3.to_csv(\"/kaggle/working/argus_sweep_snr.csv\", index=False)\n        print(f\"\\n{'SNR':>6}{'AUC s1':>9}{'AUC ver':>9}{'gain':>8}{'sd':>7}\"\n              f\"{'AUCdet s1':>11}{'AUCdet ver':>12}{'cov tgt':>9}{'cov dec':>9}\")\n        print(\"-\"*80)\n        curve = {}\n        for snr in sorted(s3.snr_db.unique(), reverse=True):\n            d  = s3[s3.snr_db == snr]\n            s1 = d[d.fusion == \"s1\"]; ve = d[d.fusion == \"ver\"]\n            # coverage is fusion-independent, so either subset reports it\n            # EVERY statistic gets its OWN noise bound. Until 28 Aug 2026 this dict computed\n            # exactly two sds -- both from `ve`, the VERIFIED subset -- and trend() was handed\n            # one of them for all four tests below. That silently mismatched three of the four:\n            #   [1]/[2] test the GAIN (ver - s1) but were bounded by the sd of `ver` ALONE;\n            #   [4] tests a STAGE-1 statistic but was bounded by the VERIFIED series' sd.\n            # Only [3] (a verified statistic bounded by verified sd) was ever scored correctly.\n            # The [4] mismatch was the damaging one: stage-1's localised AUC is ~1.7x noisier\n            # than the verifier's, so the bound was 1.7x too TIGHT and [4] printed SHRINKS for a\n            # change that does not clear stage-1's own noise. That verdict reached the public\n            # site as \"the verifier is what confers the SNR-invariance\" -- retracted 28 Aug 2026,\n            # see the lab notebook Day 34. The gain bound is paired per (species, seed) because\n            # sd(A-B) != sd(A) when A and B are correlated across the same eval pass.\n            gain_paired = (ve.sort_values([\"species\", \"seed\"]).auc.to_numpy()\n                           - s1.sort_values([\"species\", \"seed\"]).auc.to_numpy())\n            curve[int(snr)] = c = dict(\n                auc_s1=float(s1.auc.mean()), auc_ver=float(ve.auc.mean()),\n                gain=float(ve.auc.mean() - s1.auc.mean()),\n                det_s1=float(s1.auc_det.mean()), det_ver=float(ve.auc_det.mean()),\n                cov_t=float(ve.cov_target.mean()), cov_d=float(ve.cov_decoy.mean()),\n                sd=float(ve.auc.std(ddof=1)) if len(ve) > 1 else 0.0,\n                sd_gain=float(gain_paired.std(ddof=1)) if len(gain_paired) > 1 else 0.0,\n                sd_det=float(ve.auc_det.std(ddof=1)) if len(ve) > 1 else 0.0,\n                sd_det_s1=float(s1.auc_det.std(ddof=1)) if len(s1) > 1 else 0.0)\n            print(f\"{snr:>+5d}dB{c['auc_s1']:>9.3f}{c['auc_ver']:>9.3f}{c['gain']:>+8.3f}\"\n                  f\"{c['sd']:>7.3f}{c['det_s1']:>11.3f}{c['det_ver']:>12.3f}\"\n                  f\"{c['cov_t']:>9.0%}{c['cov_d']:>9.0%}\")\n\n        COV_MIN = 0.75    # below this, raw AUC is contaminated by detection failure (see caveat)\n        KNIFE   = 0.25    # a margin under 25% of the bound is a coin flip, not a finding\n\n        def trend(label, keys, get, sd_key, note, flat_note):\n            \"\"\"One trend test, reporting the MARGIN explicitly.\n\n            A bare `drop > bound` turns a 0.001 difference into a printed 'CONFIRMED' -- which\n            is how this project has produced retractions before. The margin is printed every\n            time, and the verdict is THREE-WAY, not two:\n\n              margin >  +KNIFE*bound : |change| clears the noise -> a real trend\n              margin <  -KNIFE*bound : |change| sits far BELOW the noise -> confidently FLAT\n              otherwise              : |change| ~= the noise -> genuinely unresolved\n\n            The middle and bottom cases are opposites and must not be merged. An earlier\n            version of this function tested only `margin < KNIFE*bound` and so reported a\n            change of exactly 0.000 against a 0.045 bound -- the strongest possible evidence\n            of flatness -- as \"too close to call\", burying the actual result (23 Aug 2026).\n            \"\"\"\n            if len(keys) < 2:\n                print(f\"\\n{label}\\n  -> NO VERDICT: {len(keys)} level(s) is a point, not a trend.\")\n                return\n            v_hi, v_lo = get(curve[keys[0]]), get(curve[keys[-1]])\n            if not (v_hi == v_hi and v_lo == v_lo):          # nan guard\n                print(f\"\\n{label}\\n  -> NO VERDICT: not measurable at an endpoint (nan).\")\n                return\n            drop   = v_hi - v_lo\n            bound  = max(NOISE_FLOOR, max(curve[k][sd_key] for k in keys))\n            margin = abs(drop) - bound\n            print(f\"\\n{label}\")\n            print(f\"  {v_hi:+.3f} at {keys[0]:+d} dB -> {v_lo:+.3f} at {keys[-1]:+d} dB    \"\n                  f\"change {drop:+.3f}   bound {bound:.3f}   margin {margin:+.4f}\")\n            if margin > KNIFE * bound:\n                if drop > 0:\n                    print(f\"  -> SHRINKS by {drop:.3f}, clear of the noise bound. {note}\")\n                else:\n                    print(f\"  -> GROWS by {-drop:.3f} as calls get fainter. Report as observed; do\"\n                          f\"\\n     not rationalise it into the write-up without a separate test.\")\n            elif margin < -KNIFE * bound:\n                print(f\"  -> FLAT within noise. |change| ({abs(drop):.3f}) sits well below the \"\n                      f\"noise bound\\n     ({bound:.3f}) -- this is positive evidence of no trend, \"\n                      f\"not an absent result. {flat_note}\")\n            else:\n                print(f\"  -> TOO CLOSE TO CALL. |change| and the noise bound differ by \"\n                      f\"{margin:+.4f}, within\\n     {KNIFE:.0%} of the bound itself -- a coin flip. \"\n                      f\"Unresolved: widen SNR_SWEEP_SEEDS\\n     before claiming flat OR a direction.\")\n\n        allk   = sorted(curve, reverse=True)\n        validk = sorted([k for k, c in curve.items() if c[\"cov_t\"] >= COV_MIN], reverse=True)\n\n        trend(\"[1] verification GAIN, raw AUC, full SNR range\", allk,\n              lambda c: c[\"gain\"], \"sd_gain\",\n              \"Quote the centrepiece as a high-SNR result from here on.\",\n              \"The verification gain holds across the whole tested SNR range.\")\n        if validk != allk:\n            trend(f\"[2] verification GAIN, raw AUC, coverage>={COV_MIN:.0%} levels only\", validk,\n                  lambda c: c[\"gain\"], \"sd_gain\",\n                  \"Holds where detection is still intact -- a real specificity effect.\",\n                  \"Where detection is intact, the gain does not depend on SNR.\")\n        trend(\"[3] verified AUC among LOCALISED items only (detection-corrected)\", allk,\n              lambda c: c[\"det_ver\"], \"sd_det\",\n              \"Specificity itself degrades, independently of detection.\",\n              \"Species discrimination is SNR-INVARIANT once a call is localised -- so any fall\"\n              \"\\n     in raw AUC is detection loss, not species confusion. This is the strong result.\")\n        trend(\"[4] stage-1 AUC among LOCALISED items only (control)\", allk,\n              lambda c: c[\"det_s1\"], \"sd_det_s1\",\n              \"Stage 1's own specificity degrades too.\",\n              \"Stage 1's own specificity is stable too, so [3] is not an artefact of the verifier.\")\n\n        print(f\"\\n{'-'*80}\\nHOW TO READ [1] AND [3] TOGETHER -- they answer different questions\")\n        print(\"  raw falls + detection-corrected FLAT  -> DETECTION degrades with SNR; species\")\n        print(\"                                           discrimination, given a detection, does not.\")\n        print(\"  raw falls + detection-corrected FALLS -> specificity genuinely degrades; the\")\n        print(\"                                           limitation is real and belongs in the paper.\")\n        print(\"  both flat                             -> robust across the whole tested range.\")\n        print(\"  [4] is the control: if stage-1's corrected curve moves while the verified one does\")\n        print(\"  not, the verifier is what stabilises it -- and that is the claim worth making.\")\n\n        lo = min(curve)\n        if curve[lo][\"cov_t\"] < COV_MIN:\n            bad = [k for k in allk if curve[k][\"cov_t\"] < COV_MIN]\n            print(f\"\\n  CAVEAT -- raw AUC at {bad} dB is CONTAMINATED. Target coverage there is \"\n                  f\"{curve[lo]['cov_t']:.0%};\\n  score_spec scores an unlocalised item 0.0 and _auc \"\n                  f\"counts 0.0-vs-0.0 as a tie, so those levels\\n  measure detection failure as much \"\n                  f\"as species confusion. Worse, coverage is ASYMMETRIC\\n  (target \"\n                  f\"{curve[lo]['cov_t']:.0%} vs decoy {curve[lo]['cov_d']:.0%}): unmatched zeros bias \"\n                  f\"the AUC in a fixed\\n  direction rather than merely adding noise. Use the AUCdet \"\n                  f\"columns at those levels, or\\n  restrict the claim to the coverage>={COV_MIN:.0%} \"\n                  f\"range and say so explicitly.\")\n    print(f\"[S3 took {(time.time()-t0)/60:.1f} min]\")\n\n    # ---------------------------------------------------------- S4: decoy difficulty sweep\n    # Roadmap 3.3. Decoys have always been the top-N_DECOYS nearest species in Perch space --\n    # the hardest possible confusers, chosen on purpose. That choice was never actually\n    # measured against the alternative: how much of the headline number is because the decoys\n    # are hard? This turns one number into a difficulty curve, and pre-empts \"you picked easy\n    # decoys\" -- you didn't, but until this ran that was an assertion, not a result.\n    #\n    # PRE-REGISTERED, and split into what's actually at risk of being wrong:\n    #   Not a real prediction (true by construction of how DECOYS is ranked): AUC should rise\n    #   hard -> medium -> random, for both stage-1 and verified. Decoys get progressively less\n    #   similar to the target by definition, so this ordering isn't a finding, it's a sanity\n    #   check on the ranking itself.\n    #   The actual open question: does VERIFICATION'S GAIN (ver - s1) shrink, grow, or stay\n    #   flat as decoys get easier? No strong prior either way -- reported honestly, whichever\n    #   way it lands.\n    DECOY_SWEEP_SPECIES = SWEEP_SPECIES\n    DECOY_SWEEP_SEEDS   = SWEEP_SEEDS\n    DECOY_MODES         = (\"hard\", \"medium\", \"random\")\n\n    print(f\"\\n{'='*84}\\nS4  DECOY DIFFICULTY  (roadmap 3.3)\\n{'='*84}\")\n    t0 = time.time(); rows_s4 = []\n    for mode in DECOY_MODES:\n        for sp in DECOY_SWEEP_SPECIES:\n            r = run_one(sp, N_SHOT, SHOT_STRATEGY, seeds=DECOY_SWEEP_SEEDS, decoy_mode=mode)\n            if r: rows_s4.append(r); print(f\"  {mode:<7} {sp:<12} decoy-sim {r['decoy_sim_top1']:.3f}  \"\n                                           f\"AUC {r['auc_s1']:.3f} -> {r['auc_ver']:.3f}   \"\n                                           f\"F1 {r['f1_s1']:.3f} -> {r['f1_ver']:.3f}\")\n    s4 = pd.DataFrame(rows_s4)\n\n    if len(s4):\n        s4.to_csv(\"/kaggle/working/argus_sweep_decoy_difficulty.csv\", index=False)\n        print(f\"\\n{'mode':<9}{'decoy-sim':>11}{'AUC stage-1':>13}{'AUC verified':>14}{'gain':>8}{'gain sd':>9}\")\n        print(\"-\"*64)\n        agg = {}\n        for mode, g in s4.groupby(\"decoy_mode\", sort=False):\n            gains = g.auc_ver - g.auc_s1\n            agg[mode] = dict(sim=g.decoy_sim_top1.mean(), s1=g.auc_s1.mean(), ver=g.auc_ver.mean(),\n                             gain=gains.mean(), gain_sd=gains.std(ddof=1) if len(gains) > 1 else 0.0)\n            c = agg[mode]\n            print(f\"{mode:<9}{c['sim']:>11.3f}{c['s1']:>13.3f}{c['ver']:>14.3f}{c['gain']:>+8.3f}\"\n                  f\"{c['gain_sd']:>9.3f}\")\n\n        ordered = [agg[m] for m in DECOY_MODES if m in agg]\n        print(f\"\\nsanity checks -- judged on AUC, which is what 'harder' actually MEANS:\")\n        # AUC ordering is the real difficulty check. If the tiers are separating at all, a\n        # harder decoy set must score LOWER, for stage-1 and verified alike.\n        auc_ok_s1  = all(ordered[i][\"s1\"]  <= ordered[i+1][\"s1\"]  + 1e-9 for i in range(len(ordered)-1))\n        auc_ok_ver = all(ordered[i][\"ver\"] <= ordered[i+1][\"ver\"] + 1e-9 for i in range(len(ordered)-1))\n        print(f\"  AUC rises hard -> medium -> random, stage-1:  \"\n              f\"{'OK' if auc_ok_s1 else 'VIOLATED -- tiers are not separating, investigate'}\")\n        print(f\"  AUC rises hard -> medium -> random, verified: \"\n              f\"{'OK' if auc_ok_ver else 'VIOLATED -- tiers are not separating, investigate'}\")\n        # The one similarity relation that IS true by construction: 'hard' takes ranked[0],\n        # the literal maximum, so nothing can beat it on top-1.\n        if \"hard\" in agg:\n            hard_is_max = all(agg[\"hard\"][\"sim\"] >= agg[m][\"sim\"] - 1e-9 for m in agg)\n            print(f\"  'hard' holds the highest top-1 similarity (true by construction):  \"\n                  f\"{'OK' if hard_is_max else 'VIOLATED -- decoy RANKING itself is broken'}\")\n        sim_str = \"   \".join(f\"{m}: sim {agg[m]['sim']:.3f}\" for m in DECOY_MODES if m in agg)\n        print(f\"\\n  {sim_str}\")\n        print(\"  NOTE: top-1 similarity does NOT fall hard > medium > random, and that is\")\n        print(\"  EXPECTED rather than a fault. 'medium' deliberately excludes ranks 1..N_DECOYS,\")\n        print(\"  so its top-1 is a deterministic ceiling at ranked[N_DECOYS]. 'random' samples the\")\n        print(\"  WHOLE pool -- with 30 candidates and 4 draws it catches one of the four hardest\")\n        print(\"  ~45% of the time -- and top-1 is a MAX over that draw, so it lands ABOVE medium's\")\n        print(\"  ceiling. 'random' is therefore average-difficulty-with-high-variance, not an easy\")\n        print(\"  tier. Read difficulty off the AUC columns, never off this scalar: the same lesson\")\n        print(\"  the decoy-pool confound already taught (Day 28) -- one similarity number does not\")\n        print(\"  capture a decoy SET's difficulty.\")\n\n        if \"hard\" in agg and \"random\" in agg:\n            drop = agg[\"random\"][\"gain\"] - agg[\"hard\"][\"gain\"]\n            bound = max(NOISE_FLOOR, agg[\"hard\"][\"gain_sd\"], agg[\"random\"][\"gain_sd\"])\n            margin = abs(drop) - bound\n            KNIFE = 0.25\n            print(f\"\\nTHE ACTUAL QUESTION -- does verification's gain depend on decoy difficulty?\")\n            print(f\"  gain on hardest decoys   {agg['hard']['gain']:+.3f}\")\n            print(f\"  gain on random decoys    {agg['random']['gain']:+.3f}\")\n            print(f\"  change {drop:+.3f}   bound {bound:.3f}   margin {margin:+.4f}\")\n            if margin > KNIFE * bound:\n                if drop > 0:\n                    print(f\"  -> GAIN IS LARGER against easy decoys, clear of noise. Verification's\")\n                    print(f\"     advantage is partly a property of how hard the decoys are, not just\")\n                    print(f\"     the target species -- the 0.751-class headline number is a HARD-DECOY\")\n                    print(f\"     number and should keep being described as one.\")\n                else:\n                    print(f\"  -> GAIN IS LARGER against the hardest decoys, clear of noise. Verification\")\n                    print(f\"     earns its keep precisely where it's needed most -- the headline number\")\n                    print(f\"     is not being flattered by an easy comparison; if anything it's the\")\n                    print(f\"     hardest case, understating what an easier deployment would see.\")\n            elif margin < -KNIFE * bound:\n                print(f\"  -> FLAT within noise. Verification's advantage does not depend on how hard\")\n                print(f\"     the decoys are -- the 0.751-class headline number generalises across the\")\n                print(f\"     difficulty curve, not just at the one operating point it was measured at.\")\n            else:\n                print(f\"  -> TOO CLOSE TO CALL, within {KNIFE:.0%} of the bound. Unresolved at this\")\n                print(f\"     seed count -- widen DECOY_SWEEP_SEEDS before claiming a direction.\")\n\n        print(f\"\\nWhy this matters beyond the sanity check: the headline 0.751-class AUC has always\")\n        print(f\"been reported on the HARDEST decoy set on purpose -- the number a skeptical judge\")\n        print(f\"would ask for. The random-decoy row above is the number if we'd taken the easy path\")\n        print(f\"instead, so the difference between them is now a stated fact, not an assertion.\")\n    print(f\"[S4 took {(time.time()-t0)/60:.1f} min]\")\n\nelse:\n    print(f\"\\n{'='*84}\")\n    print(\"SKIP_SWEEPS=True -- skipping N1/S2/S3/S4 (already answered twice, \"\n          \"16 Aug and 25 Aug; see the lab notebook Day 34 for the run-4 cross-check).\")\n    print(f\"{'='*84}\")\n# ------------------------------------------------------------------ N7: Perch version\n# Two-run protocol: a kernel holds ONE Perch model, so this cannot be a re-score.\n# DO THIS BEFORE THE FULL MULTI-SPECIES RUN -- decoys are ranked by Perch embedding\n# proximity, so changing Perch changes the decoy sets and makes runs incomparable.\nprint(f\"\\n{'='*84}\\nN7  PERCH 1.0 vs 2.0 -- manual two-run protocol, DO THIS FIRST\\n{'='*84}\")\nprint(f\"\"\"Currently loaded: Perch {PERCH_VERSION}, {EMBED_DIM}-d embeddings, peak-norm {PERCH_PEAK_NORM}\n\nTo compare:\n  1. Run the whole notebook with PERCH_VERSION = \"v2\" (default), SMOKE_TEST = True\n     -> save argus_multispecies_results.csv as ..._perchv2.csv\n  2. Set PERCH_VERSION = \"v1\" in CELL 1 and attach the v1 Kaggle model variation\n     (bird-vocalization-classifier/tensorFlow2/bird-vocalization-classifier, version 8)\n  3. Restart the kernel and re-run -> save as ..._perchv1.csv\n  4. Compare verified AUC per species.\n\nPublished expectation (Perch 2.0 paper, Table 3): BirdSet AUROC 0.839 -> 0.908 for the\nsame embeddings, no fine-tuning. If ARGUS sees no gain, that is worth reporting -- it\nwould suggest the bottleneck is stage 1, not the embedding.\n\nNOTE the confound: peak normalisation was ALSO wrong in v1 (raw waveforms were fed to\nPerch; the official wrapper DC-removes and scales to peak 0.25). To attribute a change to\nthe model rather than the preprocessing, run v1 twice -- PERCH_PEAK_NORM = None and 0.25 --\nor accept that the comparison bundles both fixes and say so explicitly.\"\"\")\n\n# ------------------------------------------------------------------ N2: front-end\n# NO LONGER A MANUAL PROTOCOL. CELL 2 now trains one encoder per front-end in a single\n# pass over the audio (build_banks), so PCEN vs log-Mel is a proper factor in the\n# ablation ladder (CELL 4) rather than a two-run comparison across kernels.\nprint(f\"\\n{'='*84}\\nN2  FRONT-END (PCEN vs log-Mel) -- now handled by CELL 4\\n{'='*84}\")\nprint(\"\"\"Run arguswala_ladder.py. It appears as:\n  - cumulative \"rung 4  +log-Mel\", or\n  - a main effect in the 2^4 design with FULL_FACTORIAL = True (16 configs).\n\nThe ladder resolves encoders[(FEATURE, seed)] per config, so the encoder always matches\nthe front-end -- evaluating a PCEN-trained encoder on log-Mel inputs would measure a\nmismatch, not the front-end.\n\nExpected, and genuinely contested:\n  - DCASE 2024 official table, same team and architecture: log-Mel 65.2% vs PCEN 61.5%\n  - Liang et al. Table III: log-Mel 63.67 vs PCEN 59.97 F1 -- but PCEN wins precision\n    (68.0 vs 66.4), and species specificity is a precision-shaped problem\n  - DCASE 2023 winner (Zou et al.) used PCEN at exactly ARGUS's 22050/1024/256 settings\nTwo of three point to log-Mel on F1; the 2023 winner and the precision column point back\nat PCEN. This is why it is measured rather than assumed.\"\"\")\n\nprint(f\"\\n{'='*84}\\nALL SWEEPS DONE. Copy the three CSVs out of /kaggle/working/ and log the\"\n      f\"\\nheadline numbers in the bound logbook the same day.\\n{'='*84}\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"685b10cd","cell_type":"markdown","source":"## Cell 4 — Ablation Ladder\n\n`arguswala_ladder.py`","metadata":{}},{"id":"5abd6c20","cell_type":"code","source":"# ===== CELL 4 / ABLATION LADDER =====\n# Run AFTER arguswala_v2.py (reuses its trained `encoders`, passes and scorers).\n#\n# REVISED 14 Aug 2026 after the first factorial run was killed at config 6/16.\n# Three faults in the previous version, all fixed here:\n#\n#  (1) NO CHECKPOINTING. Results were only written to CSV after the whole loop finished,\n#      so a killed session lost every completed config. ~82 min of real compute was\n#      recoverable only from scrollback. Now each config is appended to CSV the moment\n#      it finishes, and a re-run resumes instead of restarting.\n#\n#  (2) CATASTROPHIC CONFIG ORDERING. itertools.product with VERIFIER first made it the\n#      slowest-varying factor -- all 8 logreg configs ran before any proto config. The\n#      run died at 6/16, so it produced ZERO data on the prototypical probe, which was\n#      the entire reason for running a factorial. Configs are now ordered by DISTANCE\n#      FROM BASELINE, so the first 5 cover the baseline plus every single-factor main\n#      effect. A truncated run now still answers the main question.\n#\n#  (3) NO NOISE FLOOR. The previous run accidentally measured one: `pcen|logreg` scored\n#      0.696 in the cumulative run and 0.731 in the factorial run -- identical config,\n#      identical seeds, 0.035 apart, purely from GPU non-determinism (cuDNN conv backward\n#      is not bit-reproducible even with a fixed seed). Without that number, a delta of\n#      -0.005 was reported as \"HURTS (d/SE -3.02)\" when it is in fact indistinguishable\n#      from zero. The baseline is now deliberately repeated across encoder seeds to\n#      measure the floor, and every delta is judged against it.\nimport numpy as np, pandas as pd, time, itertools, os, gc\n\nLADDER_SPECIES = TARGET_LIST[:3]     # widen once the timing measurement says it fits\nLADDER_SEEDS   = (7, 8)\nFULL_FACTORIAL = True\n\n# TimeFilterAug is EXCLUDED by default as of 14 Aug 2026. Measured in isolation\n# (hn=False, both front-ends) it hurt every metric: verified AUC -0.155 (pcen) and\n# -0.088 (logmel), F1 0.713->0.522 and 0.483->0.362, and stage-1-alone AUC 0.493->0.407\n# -- that last one is verifier-independent, so the damage is unambiguously in stage-1\n# fine-tuning, not an interaction with the probe. At 4.4x the measured noise floor this\n# is a real effect, not scatter. Dropping it halves the design (16 -> 8 configs).\n# Set False to put it back and re-test.\nSKIP_TIMEFILTER = True\n\n# The baseline is re-run once per encoder seed to establish the run-to-run noise floor.\n# Deltas smaller than this are not evidence, regardless of what a within-run SE says.\nBASELINE_ENCODER_SEEDS = ENCODER_SEEDS          # from CELL 2; (0,1,2) after the widening\nOTHER_ENCODER_SEED     = ENCODER_SEEDS[0]\n\nCKPT_PATH = \"/kaggle/working/argus_ablation_ladder.csv\"\nRESUME    = True     # skip (config, encoder_seed) pairs already present in CKPT_PATH\n\n# 24 Aug 2026: the ladder finished a complete 8-config run (all 10 planned rows) that day\n# -- see the lab notebook Day 29 entry. RESUME's checkpoint only helps within ONE Kaggle\n# session (/kaggle/working/ is wiped between sessions unless manually re-attached), so a\n# fresh session re-derives an answer this project already has rather than resuming past\n# it. SKIP_LADDER=True skips this cell's run loop entirely when you only need CELL 3\n# (e.g. just the decoy-difficulty sweep, S4) and the ladder's own conclusion isn't what\n# you're re-running for. Leave False for a normal run, or when re-verifying the ladder is\n# actually the point (e.g. re-testing at a wider species count).\nSKIP_LADDER = True\n\nBASE, CAND = \"stage-1 only  (p)\", \"perch only    (v)\"\nBASELINE_CFG = dict(VERIFIER=\"logreg\", USE_TIMEFILTER=False, USE_HARD_NEG=False, FEATURE=\"pcen\")\n\nCUMULATIVE = [\n    (\"rung 0  baseline\",        dict(VERIFIER=\"logreg\", USE_TIMEFILTER=False, USE_HARD_NEG=False, FEATURE=\"pcen\")),\n    (\"rung 1  +proto probe\",    dict(VERIFIER=\"proto\",  USE_TIMEFILTER=False, USE_HARD_NEG=False, FEATURE=\"pcen\")),\n    (\"rung 2  +timefilter\",     dict(VERIFIER=\"proto\",  USE_TIMEFILTER=True,  USE_HARD_NEG=False, FEATURE=\"pcen\")),\n    (\"rung 3  +hard negatives\", dict(VERIFIER=\"proto\",  USE_TIMEFILTER=True,  USE_HARD_NEG=True,  FEATURE=\"pcen\")),\n    (\"rung 4  +log-Mel\",        dict(VERIFIER=\"proto\",  USE_TIMEFILTER=True,  USE_HARD_NEG=True,  FEATURE=\"logmel\")),\n]\n\ndef _tag(cfg):\n    t = f\"{cfg['FEATURE']}|{cfg['VERIFIER']}\"\n    if cfg[\"USE_TIMEFILTER\"]: t += \"+tfa\"\n    if cfg[\"USE_HARD_NEG\"]:   t += \"+hn\"\n    return t\n\ndef _dist_from_baseline(cfg):\n    return sum(1 for k, v in BASELINE_CFG.items() if cfg[k] != v)\n\ndef _factorial():\n    tfa_levels = (False,) if SKIP_TIMEFILTER else (False, True)\n    out = []\n    for v, t, hn, ft in itertools.product((\"logreg\", \"proto\"), tfa_levels,\n                                          (False, True), (\"pcen\", \"logmel\")):\n        cfg = dict(VERIFIER=v, USE_TIMEFILTER=t, USE_HARD_NEG=hn, FEATURE=ft)\n        out.append((_tag(cfg), cfg))\n    # Order by how many factors differ from baseline. With 4 binary factors this yields\n    # 1 baseline, then all single-factor changes, then pairs, then triples. A run killed\n    # partway still delivers every main effect -- which is what the previous ordering\n    # failed to do.\n    out.sort(key=lambda kc: (_dist_from_baseline(kc[1]), kc[0]))\n    return out\n\nCONFIGS = _factorial() if FULL_FACTORIAL else CUMULATIVE\nBASELINE_LABEL = _tag(BASELINE_CFG)\n\ndef _apply(cfg):\n    \"\"\"CELL 1 and CELL 2 share one notebook namespace, and _make_verifier(), aug_pos(),\n    stage1() and feat() all read these as globals at CALL time -- so rebinding here is\n    what switches the config. Verified by asserting the read-back below.\"\"\"\n    g = globals()\n    for k, val in cfg.items():\n        g[k] = val\n    assert (VERIFIER, USE_TIMEFILTER, USE_HARD_NEG, FEATURE) == \\\n           (cfg[\"VERIFIER\"], cfg[\"USE_TIMEFILTER\"], cfg[\"USE_HARD_NEG\"], cfg[\"FEATURE\"]), \\\n           \"flag rebind failed\"\n\ndef run_config(label, cfg, enc_seed):\n    _apply(cfg)\n    # The encoder MUST match the front-end: one trained on PCEN cannot be evaluated on\n    # log-Mel inputs, and doing so measures a mismatch rather than the front-end.\n    key = (FEATURE, enc_seed)\n    if key not in encoders:\n        raise KeyError(f\"no encoder for {key}; add '{FEATURE}' to FEATURES in CELL 2 \"\n                       f\"and re-run the encoder training. Have: {sorted(encoders)}\")\n    ENC = encoders[key]\n    rows = []\n    for sp in LADDER_SPECIES:\n        try:\n            ctx = build_target(sp, verbose=False)     # cached; only the verifier refits\n        except Exception as ex:\n            print(f\"    skip {sp}: {type(ex).__name__}: {ex}\"); continue\n        acc = {BASE: [], CAND: []}\n        for sd in LADDER_SEEDS:\n            ec = eval_pass(ENC, ctx, sd)\n            sc = spec_pass(ENC, ctx, sd)\n            nc = neg_pass(ENC, ctx, sd)\n            for name in (BASE, CAND):\n                mm = score_eval(FUSIONS[name], ec)\n                _, _, _, _, auc = score_spec(FUSIONS[name], mm[\"thr\"], sc, ctx[\"DECOYS\"])\n                acc[name].append((auc, mm[\"f1\"], mm[\"p\"], mm[\"r\"],\n                                  score_neg(FUSIONS[name], mm[\"thr\"], nc)))\n        row = dict(config=label, species=sp, encoder_seed=enc_seed, **{k: cfg[k] for k in cfg})\n        for name, tag in ((BASE, \"s1\"), (CAND, \"ver\")):\n            a = np.array(acc[name], dtype=float)\n            row[f\"auc_{tag}\"]  = a[:, 0].mean()\n            row[f\"f1_{tag}\"]   = a[:, 1].mean()\n            row[f\"prec_{tag}\"] = a[:, 2].mean()\n            row[f\"rec_{tag}\"]  = a[:, 3].mean()\n            row[f\"fa_{tag}\"]   = a[:, 4].mean()\n        rows.append(row)\n    return rows\n\n# ------------------------------------------------------------------ build the run plan\n# Baseline gets every encoder seed (noise floor); everything else gets one.\nPLAN = []\nfor label, cfg in CONFIGS:\n    seeds = BASELINE_ENCODER_SEEDS if label == BASELINE_LABEL else (OTHER_ENCODER_SEED,)\n    for es in seeds:\n        PLAN.append((label, cfg, es))\n\ndone = set()\nall_rows = []\nif RESUME and os.path.exists(CKPT_PATH):\n    prev = pd.read_csv(CKPT_PATH)\n    if {\"config\", \"encoder_seed\"}.issubset(prev.columns):\n        all_rows = prev.to_dict(\"records\")\n        done = {(r[\"config\"], int(r[\"encoder_seed\"])) for r in all_rows}\n        print(f\"RESUMING: {len(done)} (config, encoder_seed) pairs already in {CKPT_PATH}\")\n\ntodo = [(l, c, e) for (l, c, e) in PLAN if (l, int(e)) not in done]\nprint(f\"\\n{'='*100}\\nABLATION LADDER  |  {len(PLAN)} runs planned \"\n      f\"({len(CONFIGS)} configs, baseline x{len(BASELINE_ENCODER_SEEDS)} for noise floor) \"\n      f\"| {len(todo)} to run\\n\"\n      f\"  species={list(LADDER_SPECIES)}  eval seeds={LADDER_SEEDS}  \"\n      f\"timefilter={'EXCLUDED' if SKIP_TIMEFILTER else 'included'}\\n\"\n      f\"  checkpointing to {CKPT_PATH} after EVERY config -- a killed session resumes, \"\n      f\"it does not restart.\\n{'='*100}\")\n\nif SKIP_LADDER:\n    print(\"SKIP_LADDER=True -- skipping the ladder's run loop entirely.\")\n    print(\"This cell already has a complete answer from a prior session (see the\")\n    print(\"lab notebook Day 29 entry); re-running it here would just re-derive it.\")\nelse:\n    t0 = time.time()\n    for i, (label, cfg, es) in enumerate(todo, 1):\n        t1 = time.time()\n        r = run_config(label, cfg, es)\n        all_rows += r\n        # Write immediately. This is the whole point -- 82 min of compute was lost last run.\n        pd.DataFrame(all_rows).to_csv(CKPT_PATH, index=False)\n        if r:\n            d = pd.DataFrame(r)\n            print(f\"[{i:2d}/{len(todo)}] {label:<22} enc{es}  \"\n                  f\"AUC {d.auc_s1.mean():.3f} -> {d.auc_ver.mean():.3f}   \"\n                  f\"F1 {d.f1_ver.mean():.3f}   FA/h {d.fa_ver.mean():5.1f}   \"\n                  f\"[{time.time()-t1:.0f}s]  saved\")\n        gc.collect()\n        try:\n            import torch; torch.cuda.empty_cache()\n        except Exception:\n            pass\n\n    lad = pd.DataFrame(all_rows)\n    print(f\"\\ntotal {(time.time()-t0)/60:.1f} min | {len(lad)} rows -> {CKPT_PATH}\")\n\n    # ------------------------------------------------------------------ noise floor first\n    noise = None\n    b = lad[lad.config == BASELINE_LABEL]\n    if not b.empty and b.encoder_seed.nunique() > 1:\n        per_seed = b.groupby(\"encoder_seed\").auc_ver.mean()\n        noise = float(per_seed.max() - per_seed.min())\n        print(f\"\\n{'='*100}\\nNOISE FLOOR -- baseline '{BASELINE_LABEL}' repeated across \"\n              f\"{per_seed.size} encoder seeds\\n{'='*100}\")\n        print(\"  verified AUC per encoder seed: \" + \"  \".join(f\"{v:.3f}\" for v in per_seed))\n        print(f\"  spread {noise:+.3f}   sd {per_seed.std(ddof=1):.3f}\")\n        print(\"  -> any delta below this spread is NOT evidence, whatever a within-run SE says.\")\n        print(\"  (Cross-session noise is larger still: the same config scored 0.696 and 0.731 \"\n              \"in two\\n   separate sessions -- 0.035 apart -- from GPU non-determinism alone.)\")\n\n    # ------------------------------------------------------------------ report\n    print(f\"\\n{'='*100}\\n{'config':<24}{'enc':>4}{'AUC stage-1':>13}{'AUC verified':>14}\"\n          f\"{'delta vs base':>15}{'F1 ver':>9}{'FA/hr':>8}\\n{'-'*100}\")\n    base_auc = b.auc_ver.mean() if not b.empty else np.nan\n    for label, cfg in CONFIGS:\n        g = lad[lad.config == label]\n        if g.empty: continue\n        a = g.auc_ver.mean()\n        d = a - base_auc\n        flag = \"\"\n        if noise and label != BASELINE_LABEL:\n            flag = \"  <- within noise\" if abs(d) < noise else (\"  <- REAL\" if d < 0 else \"  <- REAL +\")\n        print(f\"{label:<24}{g.encoder_seed.nunique():>4}{g.auc_s1.mean():>13.3f}{a:>14.3f}\"\n              f\"{d:>+15.3f}{g.f1_ver.mean():>9.3f}{g.fa_ver.mean():>8.1f}{flag}\")\n\n    # ------------------------------------------------------------------ main effects\n    if FULL_FACTORIAL and not lad.empty:\n        print(f\"\\nMAIN EFFECTS (technique ON minus OFF, paired within species, \"\n              f\"averaged over all other settings):\")\n        for flag, on_vals, off_vals in ((\"VERIFIER\", [\"proto\"], [\"logreg\"]),\n                                        (\"USE_TIMEFILTER\", [True], [False]),\n                                        (\"USE_HARD_NEG\", [True], [False]),\n                                        (\"FEATURE\", [\"logmel\"], [\"pcen\"])):\n            if flag not in lad.columns: continue\n            g_on, g_off = lad[lad[flag].isin(on_vals)], lad[lad[flag].isin(off_vals)]\n            if g_on.empty or g_off.empty:\n                print(f\"  {flag:<34} not tested in this run\"); continue\n            # Collapse each CONFIG to one value per species BEFORE averaging configs.\n            # A plain groupby(\"species\").mean() is row-weighted, and the baseline carries\n            # 3x the rows of every other config (it is repeated across BASELINE_ENCODER_SEEDS\n            # to measure the noise floor, while challengers run at OTHER_ENCODER_SEED only).\n            # So the baseline -- the highest-scoring config -- was counted three times inside\n            # the \"off\" arm of VERIFIER, USE_HARD_NEG and FEATURE(pcen), inflating every\n            # deficit: VERIFIER read -0.048 (d/SE -5.35) when the paired value is -0.023\n            # (d/SE -1.84), which does not clear this file's own \"|d/SE| < 2 is NOT evidence\"\n            # bar. That inflated number reached the public site; retracted 28 Aug 2026, see\n            # the lab notebook Day 34. Config-weighting makes each config count once.\n            on_s  = g_on.groupby([\"species\", \"config\"]).auc_ver.mean().groupby(\"species\").mean()\n            off_s = g_off.groupby([\"species\", \"config\"]).auc_ver.mean().groupby(\"species\").mean()\n            common = on_s.index.intersection(off_s.index)\n            if len(common) < 2: continue\n            d = (on_s.loc[common] - off_s.loc[common]).to_numpy()\n            se = d.std(ddof=1) / np.sqrt(len(d)) if len(d) > 1 else float(\"nan\")\n            verdict = \"\"\n            if noise:\n                verdict = \"  within noise floor\" if abs(d.mean()) < noise else \"  exceeds noise floor\"\n            print(f\"  {flag} ({on_vals[0]} vs {off_vals[0]}):\".ljust(36)\n                  + f\"{d.mean():+.3f}  SE {se:.3f}  d/SE {d.mean()/se if se else float('nan'):>6.2f}{verdict}\")\n        print(\"  NOTE: within-run SE ignores GPU non-determinism between runs. Trust the \"\n              \"noise floor over d/SE.\")\n\n    print(\"\\nLog the noise floor and the surviving effects in the bound logbook the same day.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"98a08b27","cell_type":"markdown","source":"## Cell 5 — Confound-Closing Runs\n\n`arguswala_confounds.py`","metadata":{}},{"id":"f27bfb75","cell_type":"code","source":"# ===== CELL 5 / CONFOUND-CLOSING RUNS =====\n# Run AFTER arguswala_v2.py (reuses its trained `encoders`, passes and scorers). Does NOT\n# depend on CELL 3 or CELL 4 -- like them, it only needs CELL 1+2, so it can run in any\n# order relative to the sweeps/ladder.\n#\n# Three items the 25 Aug 2026 adversarial review named as genuinely unrun (lab notebook,\n# Day 31): \"None of these are required before Oct 3; they close INTERVIEW risk, not\n# submission risk\" -- a judge is expected to ask each of these questions live.\n#\n#  P1  PERCH ALONE AS A STANDALONE DETECTOR. Every number this project has ever reported\n#      uses stage-1 (the trained CNN) to propose candidate windows and Perch only to\n#      VERIFY them. Nobody has asked: if you skip stage-1 entirely and just slide Perch +\n#      the calibrated verifier across the recording, what do you get? Answers \"what did\n#      the CNN actually contribute\" before a judge asks it live.\n#\n#  P2  STAGE-1 FINE-TUNED WITH SPECIES-MATCHED NEGATIVES. Stage-2's Perch verifier has\n#      always trained against real other-species calls as negatives (calib_neg_species in\n#      build_target). Stage-1's per-target fine-tuning (stage1() in CELL 2) never has --\n#      its negatives have always been arbitrary random crops of whatever's in the bed.\n#      USE_SPECIES_NEG (CELL 2) tests whether that choice matters.\n#\n#  P3  RECORDIST-DISJOINT HELD-OUT SPLIT. \"Did it learn the microphone, not the bird?\"\n#      The held-out pool has never been checked against WHO recorded the 5 support shots.\n#      recordist_disjoint=True (build_target, CELL 1) additionally excludes any held-out\n#      recording sharing a recordist with a shot.\nimport numpy as np, pandas as pd, time, os, gc\n\n# 28 Aug 2026, widened after run 4: species count doubled from 3 to 5 now that\n# SKIP_MAIN_LOOP + SKIP_LADDER + SKIP_SWEEPS free up enough of the 12h ceiling to afford\n# it (budgeted ~3.8h total for a Cell-5-focused session; see the lab notebook Day 34).\nCONFOUND_SPECIES = TARGET_LIST[:5]\nCONFOUND_SEEDS   = (7, 8)\n\n# --- 29 Aug 2026: P2-only re-run, to settle the one confound that came back unresolved ---\n# The 29 Aug session (lab notebook Day 35) settled P1 (perch-alone, FALLS, clear of bound)\n# and P3 (recordist-disjoint, FLAT) at 2 seeds. P2 returned TOO CLOSE TO CALL in BOTH\n# directions, and verdict()'s own advice on that outcome is \"widen CONFOUND_SEEDS before\n# claiming a direction\" -- so that is exactly what this does, for P2 alone.\n#\n# P1/P3 stay at CONFOUND_SEEDS and are NOT re-run: their archived rows already answer their\n# questions, and re-running them at P2's seed count would triple their cost to reproduce a\n# result that is already settled. Their verdicts still print below, read from the\n# checkpoint rather than recomputed.\n# 29 Aug 2026, second revision: RUN_P2 is OFF for the distillation session. Cell 6 runs\n# AFTER this cell, so bundling them would mean an overrun kills the distillation result --\n# the one that answers the novelty question -- while preserving P2, which is a smaller\n# claim. P2 is fully set up below and costs ~3.5h whenever it is switched back on.\nRUN_P1   = False\nRUN_P2   = False\nRUN_P3   = False\nP2_SEEDS = (7, 8, 9, 10, 11, 12)   # 3x CONFOUND_SEEDS; P2 only\n\n# P2's existing rows were computed at 2 seeds and must not survive alongside 6-seed rows in\n# the same CSV -- RESUME would otherwise skip P2 entirely as \"already done\". Listing a\n# run_type here drops its rows from the checkpoint before resuming, so it recomputes clean.\nFORCE_RERUN = {\"species_neg\"}\n# 0.025 was checked against run 4, not just carried over: the configuration-matched floor\n# (same 3 species, same 2 seeds, encoder seeds 0/1/2 -- run 4's own ladder, baseline\n# verified AUC 0.754/0.739/0.756) is 0.017; per-species encoder-seed spread across the\n# CONFOUND_SPECIES themselves maxes at 0.025 (barswa). The pipeline's own run-wide 0.062\n# is driven entirely by eaywag1, which is not among CONFOUND_SPECIES -- using it here\n# would be a >2x over-correction that suppresses real effects rather than guarding\n# against noise. 0.025 is the right bound for what this file actually compares.\nNOISE_FLOOR      = 0.025\nKNIFE            = 0.25\nBASE, CAND       = \"stage-1 only  (p)\", \"perch only    (v)\"\nENC              = encoders[(FEATURE, ENCODER_SEEDS[0])]\n\n# SKIP_LADDER's _apply() (CELL 4) rebinds VERIFIER/FEATURE/USE_TIMEFILTER/USE_HARD_NEG as\n# globals and never restores them -- harmless while SKIP_LADDER=True, but if this cell is\n# ever run after a completed ladder, `ENC` above could silently bind a non-baseline\n# encoder and every build_target() here would refit a non-baseline verifier. Refuse rather\n# than produce confound numbers measured against a config nobody chose.\nassert (FEATURE, VERIFIER, USE_TIMEFILTER, USE_HARD_NEG, USE_SPECIES_NEG) == \\\n       (\"pcen\", \"logreg\", False, False, False), \\\n       \"CELL 5 inherited a non-baseline config -- check CELL 4 ran with SKIP_LADDER=True\"\n# Same guard, for CELL 6's flag. P1 in particular would be actively wrong under a distilled\n# backend: detect_perch_only() embeds with Perch by design, but ctx[\"perch_clf\"] would have\n# been fit on the student's embeddings -- scoring one space with a classifier fit in\n# another. That produces plausible numbers rather than an error, so refuse instead.\nassert EMBED_BACKEND == \"perch\", \\\n       f\"CELL 5 requires EMBED_BACKEND='perch', found {EMBED_BACKEND!r} -- CELL 6 must \" \\\n       f\"restore it (it does so in a finally block) before the confound runs are valid\"\n\nCKPT_PATH = \"/kaggle/working/argus_confounds.csv\"\nRESUME    = True     # skip (run_type, species) pairs already checkpointed\n\n# ------------------------------------------------------------------ shared plumbing\ndef _arm_metrics(fuse, ec, sc, nc, ctx):\n    m = score_eval(fuse, ec)\n    _, _, _, _, auc = score_spec(fuse, m[\"thr\"], sc, ctx[\"DECOYS\"])\n    fa = score_neg(fuse, m[\"thr\"], nc)\n    return dict(auc=auc, f1=m[\"f1\"], prec=m[\"p\"], rec=m[\"r\"], fa_per_hr=fa)\n\ndef verdict(label, val_a, val_b, sd_a, sd_b, pos_note, neg_note, flat_note,\n            noise_floor=NOISE_FLOOR, knife=KNIFE):\n    \"\"\"Two-arm version of CELL 3's trend() margin test -- same three-way logic (real /\n    flat / unresolved), just for an A-vs-B comparison instead of a multi-point curve. A\n    bare `b - a > 0` would turn a 0.001 difference into a printed verdict, which is\n    exactly the mistake CELL 3's own trend() was written to stop making.\"\"\"\n    drop = val_b - val_a\n    bound = max(noise_floor, sd_a, sd_b)\n    margin = abs(drop) - bound\n    print(f\"\\n{label}\")\n    print(f\"  {val_a:+.3f} -> {val_b:+.3f}   change {drop:+.3f}   bound {bound:.3f}   \"\n          f\"margin {margin:+.4f}\")\n    if margin > knife * bound:\n        print(f\"  -> {'RISES' if drop > 0 else 'FALLS'} by {abs(drop):.3f}, clear of the \"\n              f\"noise bound. {pos_note if drop > 0 else neg_note}\")\n    elif margin < -knife * bound:\n        print(f\"  -> FLAT within noise ({abs(drop):.3f} vs bound {bound:.3f}). {flat_note}\")\n    else:\n        print(f\"  -> TOO CLOSE TO CALL, within {knife:.0%} of the bound. Unresolved at \"\n              f\"this seed count -- widen this experiment's seed tuple (CONFOUND_SEEDS for \"\n              f\"P1/P3, P2_SEEDS for P2) before claiming a direction.\")\n\ndone = set()\nall_rows = []\nif RESUME and os.path.exists(CKPT_PATH):\n    prev = pd.read_csv(CKPT_PATH)\n    if {\"run_type\", \"species\"}.issubset(prev.columns):\n        all_rows = prev.to_dict(\"records\")\n        # Only drop what will actually be RECOMPUTED this session. Dropping a run_type\n        # whose RUN_P* flag is off would delete it from the checkpoint and never rebuild\n        # it -- silent data loss that looks like a successful run. Found 29 Aug 2026 when\n        # RUN_P2 was switched off for the distillation session while FORCE_RERUN still\n        # named \"species_neg\".\n        _RUN_FLAGS = {\"perch_alone\": RUN_P1, \"species_neg\": RUN_P2,\n                      \"recordist_disjoint\": RUN_P3}\n        _skipped = {rt for rt in FORCE_RERUN if not _RUN_FLAGS.get(rt, False)}\n        if _skipped:\n            print(f\"FORCE_RERUN names {sorted(_skipped)} but their RUN_P* flags are off \"\n                  f\"-- KEEPING their checkpointed rows rather than deleting work nothing \"\n                  f\"will rebuild.\")\n        _force = {rt for rt in FORCE_RERUN if _RUN_FLAGS.get(rt, False)}\n        if _force:\n            n_before = len(all_rows)\n            all_rows = [r for r in all_rows if r[\"run_type\"] not in _force]\n            print(f\"FORCE_RERUN {sorted(_force)}: dropped \"\n                  f\"{n_before - len(all_rows)} checkpointed rows -- they will be recomputed\")\n        done = {(r[\"run_type\"], r[\"species\"]) for r in all_rows}\n        print(f\"RESUMING: {len(done)} (run_type, species) pairs already in {CKPT_PATH}\")\n\ndef _save():\n    pd.DataFrame(all_rows).to_csv(CKPT_PATH, index=False)\n\ndef _row(run_type, species, arm, n_seeds, metrics_mean, metrics_sd, **extra):\n    return dict(run_type=run_type, species=species, arm=arm, n_seeds=n_seeds,\n                auc=metrics_mean[\"auc\"], auc_sd=metrics_sd[\"auc\"],\n                f1=metrics_mean[\"f1\"], prec=metrics_mean[\"prec\"], rec=metrics_mean[\"rec\"],\n                fa_per_hr=metrics_mean[\"fa_per_hr\"], **extra)\n\nprint(f\"\\n{'='*90}\\nCELL 5  CONFOUND-CLOSING RUNS\\n{'='*90}\")\nprint(f\"species={list(CONFOUND_SPECIES)}  seeds={CONFOUND_SEEDS}  checkpoint={CKPT_PATH}\")\nprint(f\"this session: P1={'run' if RUN_P1 else 'from checkpoint'}  \"\n      f\"P2={'run' if RUN_P2 else 'from checkpoint'} (seeds={P2_SEEDS})  \"\n      f\"P3={'run' if RUN_P3 else 'from checkpoint'}\")\n\n# ======================================================================= P1: Perch alone\n# PRE-REGISTERED: Perch was pretrained as a clip-level species classifier, never as a\n# step-by-step localiser; stage-1's CNN is specifically trained per-target, few-shot, to\n# localise. Prediction: Perch-alone, as an end-to-end detector, scores LOWER than the\n# full stage-1+verify pipeline -- possibly lower than stage-1-alone too on F1/localisation\n# -- even though Perch's species-discrimination given an already-localised candidate stays\n# strong. Consistent with this project's own AP-vs-AUC finding: Perch's strength is\n# discrimination given a candidate, not candidate generation. Reported either way.\nprint(f\"\\n{'-'*90}\\nP1  PERCH ALONE AS A STANDALONE DETECTOR\\n{'-'*90}\")\nt0 = time.time(); first_timed = False\nif not RUN_P1:\n    print(\"  RUN_P1=False -- settled 29 Aug at 2 seeds; verdict below reads the checkpoint.\")\nfor sp in (CONFOUND_SPECIES if RUN_P1 else ()):\n    if (\"perch_alone\", sp) in done:\n        print(f\"  [{sp}] already checkpointed, skipping\"); continue\n    try:\n        ctx = build_target(sp, verbose=False)\n    except Exception as ex:\n        print(f\"  skip {sp}: {type(ex).__name__}: {ex}\"); continue\n    acc = {\"s1\": [], \"ver\": [], \"perch_alone\": []}\n    for sd in CONFOUND_SEEDS:\n        ts = time.time()\n        ec_f, sc_f, nc_f = eval_pass(ENC, ctx, sd), spec_pass(ENC, ctx, sd), neg_pass(ENC, ctx, sd)\n        ec_p = eval_pass(ENC, ctx, sd, detect_fn=detect_perch_only)\n        sc_p = spec_pass(ENC, ctx, sd, detect_fn=detect_perch_only)\n        nc_p = neg_pass(ENC, ctx, sd, detect_fn=detect_perch_only)\n        acc[\"s1\"].append(_arm_metrics(FUSIONS[BASE], ec_f, sc_f, nc_f, ctx))\n        acc[\"ver\"].append(_arm_metrics(FUSIONS[CAND], ec_f, sc_f, nc_f, ctx))\n        acc[\"perch_alone\"].append(_arm_metrics(FUSIONS[CAND], ec_p, sc_p, nc_p, ctx))\n        if not first_timed:\n            first_timed = True\n            per = time.time() - ts\n            print(f\"  [timing] one (species, seed) with BOTH detectors took {per:.0f}s \"\n                  f\"-> ~{per * len(CONFOUND_SPECIES) * len(CONFOUND_SEEDS) / 60:.1f} min \"\n                  f\"projected for all of P1. Widen PERCH_ALONE_STRIDE_S (CELL 2) if that's \"\n                  f\"too slow.\")\n    for arm in acc:\n        mean = {k: np.mean([d[k] for d in acc[arm]]) for k in acc[arm][0]}\n        sd_  = {k: (np.std([d[k] for d in acc[arm]], ddof=1) if len(acc[arm]) > 1 else 0.0)\n                for k in acc[arm][0]}\n        all_rows.append(_row(\"perch_alone\", sp, arm, len(CONFOUND_SEEDS), mean, sd_))\n    _save()\n    print(f\"  [{sp}] stage1-alone AUC {all_rows[-3]['auc']:.3f}  \"\n          f\"stage1+verify AUC {all_rows[-2]['auc']:.3f}  \"\n          f\"perch-alone AUC {all_rows[-1]['auc']:.3f}  [{time.time()-t0:.0f}s elapsed]\")\n    gc.collect()\n    try:\n        import torch; torch.cuda.empty_cache()\n    except Exception:\n        pass\nprint(f\"[P1 took {(time.time()-t0)/60:.1f} min]\")\n\np1 = pd.DataFrame([r for r in all_rows if r[\"run_type\"] == \"perch_alone\"])\nif len(p1):\n    g = p1.groupby(\"arm\").agg(auc=(\"auc\", \"mean\"), auc_sd=(\"auc_sd\", \"mean\"),\n                              f1=(\"f1\", \"mean\")).reindex([\"s1\", \"ver\", \"perch_alone\"])\n    print(f\"\\n{'arm':<14}{'AUC':>8}{'F1':>8}\")\n    for arm, r in g.iterrows():\n        print(f\"{arm:<14}{r['auc']:>8.3f}{r['f1']:>8.3f}\")\n    sd_ver = float(p1[p1.arm == \"ver\"].auc_sd.mean())\n    sd_pa  = float(p1[p1.arm == \"perch_alone\"].auc_sd.mean())\n    verdict(\"Does removing stage-1 entirely change detector AUC? (verified vs perch-alone)\",\n            float(g.loc[\"ver\", \"auc\"]), float(g.loc[\"perch_alone\", \"auc\"]), sd_ver, sd_pa,\n            pos_note=\"Perch alone is BETTER than the full pipeline -- stage-1 is not \"\n                     \"adding value as a candidate proposer for these species; worth \"\n                     \"re-examining whether stage-1 is pulling its weight at all.\",\n            neg_note=\"Perch alone is WORSE than the full pipeline -- stage-1's \"\n                     \"localisation is doing real work the verifier alone cannot replace. \"\n                     \"Answers the interview question directly: the CNN's contribution is \"\n                     \"candidate generation, not species discrimination (Perch already \"\n                     \"wins that, per the AP-vs-AUC finding).\",\n            flat_note=\"Removing stage-1 doesn't move the number -- suggests Perch's own \"\n                      \"sliding-window scores are enough to localise these calls, and \"\n                      \"stage-1's contribution is smaller than assumed. Worth a wider \"\n                      \"species set before leaning on this.\")\n\n# ======================================================================= P2: species-matched negatives\n# PRE-REGISTERED: negative CHOICE is named as decisive twice already in this project's own\n# literature review (Nolasco et al. sec.4.2; Liang et al., +5.31 F1 from hard-negative\n# mining). Species-matched negatives are a strictly more informative signal than arbitrary\n# background. Prediction: stage-1-alone AUC/F1 improves with USE_SPECIES_NEG=True. No\n# strong prior on whether the gain propagates to the VERIFIED number -- the verifier\n# already handles species discrimination and may already correct for whatever stage-1\n# misses. Reported either way.\n# CAVEAT, found by audit 28 Aug 2026 before this ever ran: the \"species_matched\" arm is\n# not a clean substitution. N_NEG=200 1s calls at 0.4s gaps cannot all fit a 240s bed\n# (needs >=280s); USE_SPECIES_NEG's real placement rate is ~126/200, background-padded to\n# reach 200 total (see USE_SPECIES_NEG, CELL 2). So this tests \"~126 species-matched + ~74\n# background\" against \"~200 background\", not \"species-matched\" against \"background\" as a\n# pure factor. State results that way rather than as a clean substitution.\nprint(f\"\\n{'-'*90}\\nP2  STAGE-1 FINE-TUNED WITH SPECIES-MATCHED NEGATIVES\\n{'-'*90}\")\n\ndef _set_species_neg(flag):\n    g = globals()\n    g[\"USE_SPECIES_NEG\"] = flag\n    assert USE_SPECIES_NEG == flag, \"USE_SPECIES_NEG rebind failed\"\n\n# NOT a byte-identical A/B. eval_pass/spec_pass/neg_pass create one rng per seed and pass\n# it into detect_fn for every bed in sequence; stage1() then consumes it by an amount that\n# depends on USE_SPECIES_NEG (the True branch runs a full inject() call the False branch\n# never does), so bed 2 onward draws different planting positions between the two arms of\n# a \"same seed\" comparison. This is not new: USE_HARD_NEG in this same function already has\n# the identical property, and the whole codebase's noise-floor discipline exists because of\n# it -- CONFOUND_SEEDS averages over 2 seeds and verdict() compares against NOISE_FLOOR\n# specifically so this is a fair STATISTICAL comparison, not a claim of paired identical\n# draws. Stated explicitly here because P2 is the confound run where it matters most.\nt0 = time.time()\n_set_species_neg(False)   # defensive reset -- P1 never touches this flag, but don't assume\nif not RUN_P2:\n    print(\"  RUN_P2=False -- verdict below reads the checkpoint.\")\nfor sp in (CONFOUND_SPECIES if RUN_P2 else ()):\n    if (\"species_neg\", sp) in done:\n        print(f\"  [{sp}] already checkpointed, skipping\"); continue\n    try:\n        ctx = build_target(sp, verbose=False)\n    except Exception as ex:\n        print(f\"  skip {sp}: {type(ex).__name__}: {ex}\"); continue\n    if not ctx.get(\"neg_calls\"):\n        print(f\"  skip {sp}: no neg_calls in ctx (no non-decoy candidate species found)\")\n        continue\n    acc = {\"background|s1\": [], \"background|ver\": [],\n           \"species_matched|s1\": [], \"species_matched|ver\": []}\n    for use_species_neg, tag in ((False, \"background\"), (True, \"species_matched\")):\n        _set_species_neg(use_species_neg)\n        for sd in P2_SEEDS:\n            ec, sc, nc = eval_pass(ENC, ctx, sd), spec_pass(ENC, ctx, sd), neg_pass(ENC, ctx, sd)\n            acc[f\"{tag}|s1\"].append(_arm_metrics(FUSIONS[BASE], ec, sc, nc, ctx))\n            acc[f\"{tag}|ver\"].append(_arm_metrics(FUSIONS[CAND], ec, sc, nc, ctx))\n    _set_species_neg(False)   # reset before the next species / next confound section\n    for arm in acc:\n        mean = {k: np.mean([d[k] for d in acc[arm]]) for k in acc[arm][0]}\n        sd_  = {k: (np.std([d[k] for d in acc[arm]], ddof=1) if len(acc[arm]) > 1 else 0.0)\n                for k in acc[arm][0]}\n        all_rows.append(_row(\"species_neg\", sp, arm, len(P2_SEEDS), mean, sd_))\n    _save()\n    r = {rw[\"arm\"]: rw for rw in all_rows if rw[\"run_type\"] == \"species_neg\" and rw[\"species\"] == sp}\n    print(f\"  [{sp}] s1 AUC {r['background|s1']['auc']:.3f} -> {r['species_matched|s1']['auc']:.3f}   \"\n          f\"ver AUC {r['background|ver']['auc']:.3f} -> {r['species_matched|ver']['auc']:.3f}\")\n    gc.collect()\n    try:\n        import torch; torch.cuda.empty_cache()\n    except Exception:\n        pass\nprint(f\"[P2 took {(time.time()-t0)/60:.1f} min]\")\n\np2 = pd.DataFrame([r for r in all_rows if r[\"run_type\"] == \"species_neg\"])\nif len(p2):\n    for tag in (\"s1\", \"ver\"):\n        bg = p2[p2.arm == f\"background|{tag}\"]\n        sm = p2[p2.arm == f\"species_matched|{tag}\"]\n        if bg.empty or sm.empty: continue\n        verdict(f\"Does a species-matched negative pool change stage-1 fine-tuning? \"\n                f\"({'stage-1 alone' if tag == 's1' else 'verified'} AUC)\",\n                float(bg.auc.mean()), float(sm.auc.mean()),\n                float(bg.auc_sd.mean()), float(sm.auc_sd.mean()),\n                pos_note=\"Species-matched negatives HELP -- stage-1's fine-tuning negative \"\n                         \"choice matters, matching Liang et al./Nolasco et al. Worth \"\n                         \"adopting as a rung, not just a confound check.\",\n                neg_note=\"Species-matched negatives HURT -- plausible if the pool is small \"\n                         \"or dominated by acoustically-distant species, giving stage-1 \"\n                         \"less varied negative exposure than random background did.\",\n                flat_note=\"Negative source doesn't move stage-1 -- background windows were \"\n                          \"already an adequate negative signal for this task; the \"\n                          \"literature's finding may not transfer to this planting protocol.\")\n\n# ======================================================================= P3: recordist-disjoint\n# PRE-REGISTERED: if ARGUS's specificity number is genuinely about the SPECIES rather than\n# the recording SETUP, a recordist-disjoint held-out split should barely move AUC. A real\n# drop would mean part of the measured signal has been \"which microphone/recordist\", not\n# \"which bird\" -- a genuine confound this project has not yet ruled out. No strong prior\n# on magnitude either way; reported honestly regardless of outcome, per this project's\n# established practice throughout this notebook.\nprint(f\"\\n{'-'*90}\\nP3  RECORDIST-DISJOINT HELD-OUT SPLIT\\n{'-'*90}\")\nt0 = time.time()\nif not RUN_P3:\n    print(\"  RUN_P3=False -- settled 29 Aug at 2 seeds; verdict below reads the checkpoint.\")\nfor sp in (CONFOUND_SPECIES if RUN_P3 else ()):\n    if (\"recordist_disjoint\", sp) in done:\n        print(f\"  [{sp}] already checkpointed, skipping\"); continue\n    acc = {\"shared_ok|s1\": [], \"shared_ok|ver\": [],\n           \"disjoint|s1\": [], \"disjoint|ver\": []}\n    n_heldout, n_recordists = {}, {}\n    skip_species = False\n    for disjoint, tag in ((False, \"shared_ok\"), (True, \"disjoint\")):\n        try:\n            ctx = build_target(sp, recordist_disjoint=disjoint, verbose=False)\n        except (ValueError, KeyError) as ex:\n            print(f\"  [{sp}] {tag}: {type(ex).__name__}: {ex}\")\n            if disjoint:\n                print(f\"  skip {sp}: recordist-disjoint split leaves no held-out pool \"\n                      f\"(too few distinct recordists for this species) -- itself worth \"\n                      f\"logging, not a bug to route around.\")\n            skip_species = True\n            break\n        n_heldout[tag] = ctx[\"covariates\"][\"n_heldout_calls\"]\n        n_recordists[tag] = ctx[\"covariates\"][\"n_distinct_recordists\"]\n        for sd in CONFOUND_SEEDS:\n            ec, sc, nc = eval_pass(ENC, ctx, sd), spec_pass(ENC, ctx, sd), neg_pass(ENC, ctx, sd)\n            acc[f\"{tag}|s1\"].append(_arm_metrics(FUSIONS[BASE], ec, sc, nc, ctx))\n            acc[f\"{tag}|ver\"].append(_arm_metrics(FUSIONS[CAND], ec, sc, nc, ctx))\n    if skip_species:\n        continue\n    for arm in acc:\n        mean = {k: np.mean([d[k] for d in acc[arm]]) for k in acc[arm][0]}\n        sd_  = {k: (np.std([d[k] for d in acc[arm]], ddof=1) if len(acc[arm]) > 1 else 0.0)\n                for k in acc[arm][0]}\n        tag = arm.split(\"|\")[0]\n        all_rows.append(_row(\"recordist_disjoint\", sp, arm, len(CONFOUND_SEEDS), mean, sd_,\n                             n_heldout_calls=n_heldout[tag], n_distinct_recordists=n_recordists[tag]))\n    _save()\n    r = {rw[\"arm\"]: rw for rw in all_rows\n         if rw[\"run_type\"] == \"recordist_disjoint\" and rw[\"species\"] == sp}\n    print(f\"  [{sp}] shared-ok ver AUC {r['shared_ok|ver']['auc']:.3f} \"\n          f\"(n_heldout={n_heldout['shared_ok']})  ->  \"\n          f\"disjoint ver AUC {r['disjoint|ver']['auc']:.3f} \"\n          f\"(n_heldout={n_heldout['disjoint']}, distinct_recordists={n_recordists['disjoint']})\")\n    gc.collect()\n    try:\n        import torch; torch.cuda.empty_cache()\n    except Exception:\n        pass\nprint(f\"[P3 took {(time.time()-t0)/60:.1f} min]\")\n\np3 = pd.DataFrame([r for r in all_rows if r[\"run_type\"] == \"recordist_disjoint\"])\nif len(p3):\n    for tag in (\"s1\", \"ver\"):\n        so = p3[p3.arm == f\"shared_ok|{tag}\"]\n        dj = p3[p3.arm == f\"disjoint|{tag}\"]\n        if so.empty or dj.empty: continue\n        verdict(f\"Does a recordist-disjoint held-out split change AUC? \"\n                f\"({'stage-1 alone' if tag == 's1' else 'verified'})\",\n                float(so.auc.mean()), float(dj.auc.mean()),\n                float(so.auc_sd.mean()), float(dj.auc_sd.mean()),\n                pos_note=\"AUC is HIGHER once recordist overlap is removed -- unexpected; \"\n                         \"check the held-out pool sizes above before trusting this, small \"\n                         \"pools swing more.\",\n                neg_note=\"AUC FALLS once recordist overlap is removed -- part of the \"\n                         \"measured signal was 'recognise this recordist's setup', not \"\n                         \"purely the species. A real, reportable limitation -- state the \"\n                         \"size of the drop, don't just flag its direction.\",\n                flat_note=\"AUC does not depend on recordist overlap -- the specificity \"\n                          \"result is not an artefact of recording-setup leakage. Rules \"\n                          \"out the 'did it learn the microphone, not the bird' concern \"\n                          \"for these species.\")\n\nprint(f\"\\n{'='*90}\\nALL CONFOUND-CLOSING RUNS DONE -- {len(all_rows)} rows -> {CKPT_PATH}\")\nprint(f\"Log the three verdicts above in the bound logbook the same day.\\n{'='*90}\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"bf2ba4d1","cell_type":"markdown","source":"## Cell 6 — Perch Distillation\n\n`arguswala_distill.py`","metadata":{}},{"id":"7aa6dd9b","cell_type":"code","source":"# ===== CELL 6 / PERCH DISTILLATION =====\n# Run AFTER arguswala_v2.py (needs its `encoders`, `Enc`, `feat`/`crop`, and the three\n# evaluation passes). Independent of CELLS 3-5.\n#\n# THE QUESTION. Every species-discrimination number this project reports comes out of\n# Perch v2 -- a ~12M-parameter model Google trained on 14,795 classes. Stage 1, the part\n# built here, is a ~445k-parameter trunk (642k once the projection head needed to actually\n# stand in for Perch is included) trained on 40 species, and CELL 5's P1 established\n# that neither stage does the other's job: Perch alone collapses to 0.526 AUC / 2.5%\n# recall as a detector, stage-1 alone sits at 0.585 on discrimination. The pipeline works;\n# but the DISCRIMINATION half of it is Perch's.\n#\n# That is also a deployment problem, not only an authorship one. Perch cannot run on a\n# solar-powered recorder in a shola forest. A 642k-parameter CNN can. So:\n#\n#   Can Perch's species knowledge be compressed ~19x into the stage-1 encoder, and what\n#   does that cost in species discrimination?\n#\n# METHOD. Standard knowledge distillation. The student is the SAME Enc architecture stage 1\n# already uses, plus a linear head projecting its 128-d output to Perch's 1536-d space. It\n# is trained to reproduce Perch's embedding of the same 1-second audio, by cosine loss.\n# Perch is the TEACHER, used at training time only. At detection time under\n# EMBED_BACKEND=\"distilled\", no Perch inference runs at all.\n#\n# WHAT IS HELD FIXED, AND WHY IT MATTERS. Only the verifier's embedding space changes.\n# Decoy ranking and shot selection stay on Perch under both arms (see EMBED_BACKEND in\n# CELL 1) -- those decide WHAT THE TEST IS. Changing them alongside the thing being\n# measured is the exact error the literature review flagged for the v1->v2 upgrade. Same\n# species, same shots, same decoys, same beds, same seeds: only the embedding space moves.\n#\n# PRE-REGISTERED PREDICTION, written before the first run. The student will lose to the\n# teacher -- a 19x smaller model trained on 40 species cannot carry all of a 14,795-class\n# model's discriminative structure. The open question is HOW MUCH is lost, and the honest\n# outcomes are: (a) small loss, inside the noise floor -> compression is nearly free and\n# field deployment is real; (b) large but bounded loss -> the accuracy/deployability\n# frontier gets a number nobody has published; (c) collapse to chance -> distillation does\n# not transfer at this scale, reported as a negative result. All three are publishable.\n# What is NOT acceptable is quietly reporting only whichever the run happens to produce.\nimport numpy as np, pandas as pd, torch, torch.nn as nn, torch.optim as optim, time, os, gc\n\nSKIP_DISTILL   = True\nDISTILL_SPECIES_N   = 60      # bank species for the teacher-student set (targets excluded)\nDISTILL_RECS_PER_SP = 12      # recordings per species\nDISTILL_CROPS_PER_REC = 4     # peak windows per recording -> ~2,880 pairs at the defaults\nDISTILL_EPOCHS = 60\nDISTILL_BATCH  = 64\nDISTILL_LR     = 1e-3\nDISTILL_VAL_FRAC = 0.15       # held-out species, not held-out crops -- see below\nDISTILL_SEED   = 0\nDISTILL_EVAL_SPECIES = TARGET_LIST[:5]\nDISTILL_EVAL_SEEDS   = (7, 8)\nENC_D  = encoders[(FEATURE, ENCODER_SEEDS[0])]   # stage 1 is IDENTICAL in both arms\nCAND_D = \"perch only    (v)\"                     # the verifier-only fusion, both arms\nDISTILL_CKPT   = \"/kaggle/working/argus_distill.csv\"\n# Same bound CELL 5 uses and for the same reason: it is the configuration-matched\n# encoder-seed spread over these species, not the pipeline's run-wide figure.\nDISTILL_NOISE_FLOOR = 0.025\n\nassert EMBED_BACKEND == \"perch\", \\\n    f\"CELL 6 must start from the Perch backend, found {EMBED_BACKEND!r}\"\n\n# ------------------------------------------------------------------ teacher-student data\ndef _peak_offsets(y, k, n=CLEN):\n    \"\"\"Same selection rule as CELL 1's peak_windows(), but returns the OFFSETS rather than\n    the audio. Needed because the teacher and the student must be handed the same second:\n    peak_windows() returns only the waveform, so pairing it with a spectrogram crop would\n    mean re-deriving where it came from. Returning offsets makes the correspondence exact\n    by construction instead of by reconstruction.\"\"\"\n    if len(y) <= n: return [0]\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    w = cs[n:] - cs[:-n]\n    order = np.argsort(w)[::-1]\n    taken = []\n    for _ in range(k):\n        pick = next((int(i) for i in order if all(abs(int(i) - t) >= n for t in taken)), None)\n        if pick is None: break\n        taken.append(pick)\n    return taken\n\ndef build_distill_set(exclude, feature=FEATURE, n_species=DISTILL_SPECIES_N,\n                      recs_per_sp=DISTILL_RECS_PER_SP, crops_per_rec=DISTILL_CROPS_PER_REC,\n                      seed=DISTILL_SEED):\n    \"\"\"-> (specs, embs, species) : student inputs, teacher targets, and the species each\n    pair came from.\n\n    Deliberately a SEPARATE builder from CELL 2's build_banks(), not a modification of it.\n    build_banks() feeds the stage-1 encoder that produced every archived result in this\n    project; changing it would invalidate the centrepiece. This function reads the same\n    dataset with the same exclusions and the same front-end, but keeps the raw waveform\n    crop alongside the spectrogram crop, because the teacher needs audio and the student\n    needs a spectrogram -- of the SAME one second, which is the whole point.\n\n    crops_per_rec exists because ~480 pairs (40 species x 12 recordings) is far too few to\n    fit a 128->1536 projection, which alone has ~197k parameters. Taking the top-k loudest\n    non-overlapping windows per recording multiplies the pair count without touching a new\n    file, and peak_windows() is already the project's way of finding where a bird actually\n    calls in a mostly-silent recording.\"\"\"\n    rng = np.random.default_rng(seed)\n    sps = [s for s in meta.primary_label.unique() if s not in exclude]\n    rng.shuffle(sps)\n    specs, waves, species = [], [], []\n    picked = 0\n    for sp in sps:\n        if picked >= n_species: break\n        got = 0\n        for fn in meta[meta.primary_label == sp][\"filename\"].tolist()[:recs_per_sp]:\n            try:\n                y = ld(os.path.join(AUD, fn), 15)\n                # Featurise the WHOLE recording then crop, exactly as build_banks() and\n                # stage1()'s scan do. Featurising an isolated 1 s clip instead would give\n                # different mel edge behaviour from what the student meets at detection\n                # time -- a train/test mismatch introduced by the data builder.\n                f = feat(y, feature)\n                for off in _peak_offsets(y, k=crops_per_rec, n=CLEN):\n                    w = y[off:off + CLEN]\n                    if len(w) < CLEN: continue\n                    specs.append(crop(f, off // HOP + SEG_FR // 2))\n                    waves.append(w.astype(\"float32\"))\n                    species.append(sp)\n                    got += 1\n            except Exception:\n                pass\n        if got >= 8:\n            picked += 1\n    if not specs:\n        raise RuntimeError(\"distillation set is empty -- check AUD/meta are loaded\")\n    specs = np.stack(specs)\n    print(f\"  distill set: {len(specs)} (spectrogram, waveform) pairs \"\n          f\"from {picked} species\")\n    t0 = time.time()\n    embs = []\n    for i in range(0, len(waves), 256):\n        embs.append(perch_embed(waves[i:i + 256]))\n        if i == 0:\n            print(f\"  [timing] first 256 teacher embeddings took {time.time()-t0:.0f}s \"\n                  f\"-> ~{(time.time()-t0)*len(waves)/256/60:.1f} min projected\")\n    embs = l2n(np.concatenate(embs))\n    return specs, embs.astype(\"float32\"), np.array(species)\n\n\nclass DistillStudent(nn.Module):\n    \"\"\"The stage-1 encoder, plus a linear projection into Perch's embedding width.\n\n    The trunk is Enc() unchanged -- the same 4-conv-block, ~445k-parameter network stage 1\n    already runs -- so a positive result here is a statement about the model this project\n    actually deploys, not about some larger network introduced for the occasion. Only the\n    projection head is new, and it is deliberately linear: if a single matrix can carry\n    Perch's structure, that is a stronger and more interpretable claim than if an MLP is\n    needed to rescue it.\"\"\"\n    def __init__(s, out_dim):\n        super().__init__()\n        s.trunk = Enc()\n        s.head  = nn.Linear(EMB, out_dim)\n    def forward(s, x):\n        return s.head(s.trunk(x))\n    def embed(s, x):\n        return s.head(s.trunk(x))\n\n\ndef train_distill(specs, embs, species, epochs=DISTILL_EPOCHS, batch=DISTILL_BATCH,\n                  lr=DISTILL_LR, val_frac=DISTILL_VAL_FRAC, seed=DISTILL_SEED):\n    \"\"\"Cosine-loss distillation, validated on HELD-OUT SPECIES.\n\n    The validation split is by species, not by crop. Splitting crops at random would put\n    other windows from the same recording -- often the same individual bird, the same\n    microphone, the same afternoon -- on both sides, and the reported validation number\n    would then measure memorisation rather than transfer. Since the entire point is\n    whether the student generalises to species it has never met (which is what a five-shot\n    target IS), the split has to be at the species level to mean anything.\n\n    Cosine rather than MSE because every downstream consumer of these embeddings\n    L2-normalises first (l2n in CELL 1, applied in build_target and detect_pv), so only\n    direction is ever used. Training to match magnitudes would spend capacity on a\n    quantity nothing reads.\"\"\"\n    rng = np.random.default_rng(seed)\n    uniq = np.unique(species); rng.shuffle(uniq)\n    n_val = max(1, int(round(len(uniq) * val_frac)))\n    val_sp = set(uniq[:n_val].tolist())\n    vm = np.array([s in val_sp for s in species])\n    print(f\"  split: {(~vm).sum()} train pairs / {vm.sum()} val pairs \"\n          f\"({len(uniq)-n_val} train species / {n_val} held-out species)\")\n\n    Xtr = torch.tensor(specs[~vm]); Ytr = torch.tensor(embs[~vm])\n    Xva = torch.tensor(specs[vm]).to(device); Yva = torch.tensor(embs[vm]).to(device)\n\n    m = DistillStudent(embs.shape[1]).to(device)\n    opt = optim.Adam(m.parameters(), lr)\n    sched = optim.lr_scheduler.CosineAnnealingLR(opt, T_max=epochs)\n    n = len(Xtr); best_val = -1.0; best_state = None\n    for ep in range(epochs):\n        m.train()\n        perm = torch.randperm(n)\n        tot = 0.0\n        for i in range(0, n, batch):\n            idx = perm[i:i + batch]\n            xb = Xtr[idx].to(device); yb = Ytr[idx].to(device)\n            pred = m(xb)\n            loss = (1 - nn.functional.cosine_similarity(pred, yb, dim=1)).mean()\n            opt.zero_grad(); loss.backward(); opt.step()\n            tot += float(loss) * len(idx)\n        sched.step()\n        m.eval()\n        with torch.no_grad():\n            vsim = float(nn.functional.cosine_similarity(m(Xva), Yva, dim=1).mean())\n        if vsim > best_val:\n            best_val = vsim\n            best_state = {k: v.detach().clone() for k, v in m.state_dict().items()}\n        if (ep + 1) % 10 == 0 or ep == 0:\n            print(f\"    ep {ep+1:3d}/{epochs}  train cos {1-tot/n:.4f}  \"\n                  f\"held-out-species cos {vsim:.4f}\")\n    # Select on held-out species rather than taking the last epoch: the last epoch is\n    # whatever the schedule happened to end on, and reporting it would confound \"how well\n    # this distils\" with \"where training stopped\".\n    m.load_state_dict(best_state)\n    m.eval()\n    print(f\"  best held-out-species cosine similarity to Perch: {best_val:.4f}\")\n    return m, best_val\n\n\ndef make_distill_embedder(model, feature=FEATURE):\n    \"\"\"-> a function with verifier_embed_ctx's contract: (y22, centres_sec) -> (n, D).\n\n    Mirrors perch_embed_ctx's geometry at the student's own scale. Perch cuts a 5-second\n    window around each event because that is its input length; the student cuts SEG_SEC\n    (1 s), because that is what it was distilled on and what stage 1 scans with. Feeding\n    it a 5-second window would be a train/test mismatch, not a fairer comparison.\"\"\"\n    def _embed_ctx(y22, centres_sec):\n        if len(centres_sec) == 0:\n            return np.zeros((0, model.head.out_features), dtype=\"float32\")\n        f = feat(y22, feature)\n        xs = [crop(f, int(round(c * SR / HOP))) for c in centres_sec]\n        out = []\n        with torch.no_grad():\n            for i in range(0, len(xs), 256):\n                b = torch.tensor(np.stack(xs[i:i + 256])).to(device)\n                out.append(model.embed(b).cpu().numpy())\n        return np.concatenate(out).astype(\"float32\")\n    return _embed_ctx\n\n\n# ------------------------------------------------------------------ run\nrows = []\nif SKIP_DISTILL:\n    print(\"SKIP_DISTILL=True -- skipping CELL 6 entirely.\")\nelse:\n    print(f\"\\n{'='*90}\\nCELL 6  PERCH DISTILLATION\\n{'='*90}\")\n    print(f\"teacher=Perch {PERCH_VERSION} ({EMBED_DIM}-d)   \"\n          f\"student=Enc + Linear({EMB}->{EMBED_DIM})\")\n    t0 = time.time()\n    specs, embs, species = build_distill_set(TARGET_SET)\n    print(f\"[teacher embeddings took {(time.time()-t0)/60:.1f} min]\")\n\n    t0 = time.time()\n    student, val_cos = train_distill(specs, embs, species)\n    print(f\"[distillation took {(time.time()-t0)/60:.1f} min]\")\n\n    n_student = sum(p.numel() for p in student.parameters())\n    print(f\"\\nstudent parameters: {n_student:,}  \"\n          f\"(trunk {sum(p.numel() for p in student.trunk.parameters()):,} + \"\n          f\"head {sum(p.numel() for p in student.head.parameters()):,})\")\n\n    # ---------------------------------------------------------- teacher vs student, A/B\n    # Same species, same shots, same decoys, same beds, same seeds. The ONLY difference\n    # between the two arms is which embedding space the verifier lives in.\n    print(f\"\\n{'-'*90}\\nTEACHER vs STUDENT (same test, only the embedding space differs)\"\n          f\"\\n{'-'*90}\")\n    DISTILL_EMBED_CTX = make_distill_embedder(student)\n    globals()[\"DISTILL_EMBED_CTX\"] = DISTILL_EMBED_CTX\n\n    acc = {\"perch\": [], \"distilled\": []}\n    try:\n        for backend in (\"perch\", \"distilled\"):\n            globals()[\"EMBED_BACKEND\"] = backend\n            assert EMBED_BACKEND == backend, \"EMBED_BACKEND rebind failed\"\n            for sp in DISTILL_EVAL_SPECIES:\n                try:\n                    ctx = build_target(sp, verbose=False)\n                except Exception as ex:\n                    print(f\"  skip {sp} ({backend}): {type(ex).__name__}: {ex}\"); continue\n                per_seed = []\n                for sd in DISTILL_EVAL_SEEDS:\n                    ec = eval_pass(ENC_D, ctx, sd)\n                    sc = spec_pass(ENC_D, ctx, sd)\n                    nc = neg_pass(ENC_D, ctx, sd)\n                    m = score_eval(FUSIONS[CAND_D], ec)\n                    _, _, _, _, auc = score_spec(FUSIONS[CAND_D], m[\"thr\"], sc, ctx[\"DECOYS\"])\n                    per_seed.append(dict(auc=auc, f1=m[\"f1\"], prec=m[\"p\"], rec=m[\"r\"],\n                                         fa_per_hr=score_neg(FUSIONS[CAND_D], m[\"thr\"], nc)))\n                mean = {k: float(np.mean([d[k] for d in per_seed])) for k in per_seed[0]}\n                sd_  = {k: float(np.std([d[k] for d in per_seed], ddof=1))\n                        if len(per_seed) > 1 else 0.0 for k in per_seed[0]}\n                acc[backend].append(mean)\n                rows.append(dict(backend=backend, species=sp,\n                                 n_seeds=len(DISTILL_EVAL_SEEDS),\n                                 student_params=n_student, val_cosine=val_cos,\n                                 auc=mean[\"auc\"], auc_sd=sd_[\"auc\"], f1=mean[\"f1\"],\n                                 prec=mean[\"prec\"], rec=mean[\"rec\"],\n                                 fa_per_hr=mean[\"fa_per_hr\"]))\n                pd.DataFrame(rows).to_csv(DISTILL_CKPT, index=False)\n                print(f\"  [{backend:9s} {sp}] AUC {mean['auc']:.3f}  F1 {mean['f1']:.3f}  \"\n                      f\"FA/hr {mean['fa_per_hr']:.1f}\")\n                gc.collect()\n                try: torch.cuda.empty_cache()\n                except Exception: pass\n    finally:\n        # Restore unconditionally. CELL 5 asserts EMBED_BACKEND==\"perch\" precisely because\n        # leaving it flipped would make every later cell score a classifier in the wrong\n        # space -- silently, with plausible numbers.\n        globals()[\"EMBED_BACKEND\"] = \"perch\"\n        print(f\"\\n  EMBED_BACKEND restored to {EMBED_BACKEND!r}\")\n\n    # ---------------------------------------------------------- verdict\n    if acc[\"perch\"] and acc[\"distilled\"]:\n        pa = float(np.mean([d[\"auc\"] for d in acc[\"perch\"]]))\n        da = float(np.mean([d[\"auc\"] for d in acc[\"distilled\"]]))\n        df = pd.DataFrame(rows)\n        sd_p = float(df[df.backend == \"perch\"].auc_sd.mean())\n        sd_d = float(df[df.backend == \"distilled\"].auc_sd.mean())\n        drop = da - pa\n        bound = max(DISTILL_NOISE_FLOOR, sd_p, sd_d)\n        print(f\"\\n{'='*90}\")\n        print(f\"Perch verifier      mean species AUC  {pa:.4f}\")\n        print(f\"Distilled verifier  mean species AUC  {da:.4f}\")\n        print(f\"change {drop:+.4f}   bound {bound:.4f}   \"\n              f\"student is {12_000_000/n_student:.0f}x smaller than Perch\")\n        if abs(drop) - bound < -0.25 * bound:\n            print(\"-> COMPRESSION IS ESSENTIALLY FREE at this scale: the student matches \"\n                  \"the teacher within the measured noise bound, with no Perch inference \"\n                  \"at detection time. This is the deployable result.\")\n        elif drop < 0:\n            print(f\"-> The student LOSES {abs(drop):.3f} AUC. Report it as the measured \"\n                  f\"cost of {12_000_000/n_student:.0f}x compression -- an \"\n                  f\"accuracy/deployability frontier, not a failure, provided it stays \"\n                  f\"clear of chance (0.50).\")\n        else:\n            print(\"-> The student MATCHES OR BEATS the teacher. Check for leakage before \"\n                  \"believing it: confirm the validation split was by species, and that \"\n                  \"decoy ranking still ran on Perch under both arms.\")\n        print(f\"{'='*90}\")\n        print(f\"wrote {DISTILL_CKPT}  ({len(rows)} rows)\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"0e7f3671","cell_type":"markdown","source":"## Cell 7 — Bottleneck Frontier & Wild-Audio Scan\n\n`arguswala_frontier.py`","metadata":{}},{"id":"859ebad8","cell_type":"code","source":"# ===== CELL 7 / BOTTLENECK FRONTIER + WILD-AUDIO SCAN =====\n# Run AFTER arguswala_v2.py and arguswala_distill.py (reuses CELL 6's build_distill_set and\n# make_distill_embedder, which are defined even when SKIP_DISTILL=True). Independent of\n# CELLS 3-5.\n#\n# Two experiments, chosen because each attacks a limitation the research paper currently\n# states outright, and because neither can be answered by more of the same compute.\n#\n# ---------------------------------------------------------------------------------------\n# PART A -- THE BOTTLENECK FRONTIER. Why the distilled verifier failed, and whether it has\n# to.\n#\n# CELL 6 distilled Perch into the stage-1 network and lost 0.136 mean species AUC, with\n# S. albiventris collapsing to 0.511 (chance). That was reported honestly as bounding \"a\n# modest first attempt\", not distillation in general. This tests the most specific\n# mechanical explanation available for why it failed:\n#\n#   Perch emits 1536 dimensions. The student's trunk emits 128. A linear head from 128-d\n#   can only ever reproduce a 128-dimensional SUBSPACE of Perch's space, no matter how\n#   long it trains or how much data it sees. If Perch's species-discriminative structure\n#   needs more than 128 dimensions, the student is not undertrained -- it is architecturally\n#   incapable, and CELL 6's held-out-species cosine of 0.399 is a ceiling, not a symptom.\n#\n# So: sweep the bottleneck width and watch what happens to distillation fidelity AND to\n# downstream species AUC. This produces an accuracy-vs-size curve rather than a single\n# pass/fail, which is what the paper's own framing (\"an accuracy/deployability frontier\")\n# actually promised but did not yet deliver.\n#\n# PRE-REGISTERED, before running. If the bottleneck is the binding constraint, held-out\n# cosine and species AUC both rise monotonically with width, and the curve says how much\n# model a field deployment has to carry to keep how much discrimination -- a genuinely\n# useful number nobody has published. If both stay flat, the bottleneck is NOT the\n# constraint, CELL 6's failure is about data or optimisation instead, and that redirects\n# the future work in the paper. A third outcome is possible and must not be hidden: fidelity\n# (cosine) rises while species AUC does not, which would mean the student is reproducing\n# Perch's embedding geometry in aggregate while losing exactly the part that separates\n# confusable species -- the most interesting outcome of the three, and the one most relevant\n# to this project's own thesis about metrics measuring the wrong thing.\n#\n# A nonlinear-head arm at the widest setting separates \"not enough dimensions\" from \"not\n# enough expressiveness in the map\". Same trunk, same data, same everything else.\n#\n# ---------------------------------------------------------------------------------------\n# PART B -- WILD-AUDIO SCAN. The pipeline has never been asked to find a real call.\n#\n# Limitation #2 in the paper: every reported number comes from target calls PLANTED into\n# real soundscape at controlled SNR. The system has never once been pointed at continuous\n# unlabelled forest audio and asked what is in it. That is the gap between a benchmark and\n# an instrument.\n#\n# This scans soundscapes the pipeline has never touched -- strictly disjoint from the 14\n# used for calibration and evaluation -- at the same best-F1 threshold every other number\n# in this project uses, and saves the highest-scoring detections as listenable clips.\n#\n# WHAT THIS CAN AND CANNOT ESTABLISH, stated before it runs. It cannot confirm a detection:\n# nobody in this project can reliably identify S. albiventris by ear, and there is no ground\n# truth for these files. It CAN establish a detection RATE on genuinely wild audio, produce\n# a ranked shortlist short enough for a human (or later, an expert) to adjudicate, and\n# compare that rate against the same pipeline hunting a different species in the identical\n# audio. If the study species and the control fire at indistinguishable rates, that is\n# evidence the detector is responding to generic bird-shaped sound rather than to this\n# species, and it must be reported that way.\nimport numpy as np, pandas as pd, torch, torch.nn as nn, torch.optim as optim\nimport time, os, gc, glob\n\nSKIP_FRONTIER = False\n\n# ------------------------------------------------------------------ Part A config\nFRONTIER_WIDTHS      = (64, 128, 256, 512)\nFRONTIER_NONLINEAR   = True          # extra arm: MLP head at the widest width\nFRONTIER_EVAL_SPECIES = TARGET_LIST[:5]\nFRONTIER_EVAL_SEEDS   = (7, 8)\nRUN_PERCH_CONTROL     = True         # same-session Perch baseline, not a carried-over number\n\n# Distillation set, scaled up from CELL 6's 2,636 pairs / 60 species. Species diversity is\n# raised in preference to crops-per-recording: CELL 6 hit train cosine 0.669 against\n# held-out-species 0.399, and a train/held-out gap that wide is a generalisation problem\n# across species, which more crops of the SAME species does not fix.\nFRONTIER_SPECIES_N   = 98            # effectively every non-target species available\nFRONTIER_RECS_PER_SP = 20\nFRONTIER_CROPS_PER_REC = 6\nFRONTIER_EPOCHS      = 150\nFRONTIER_BATCH       = 128\nFRONTIER_LR          = 1e-3\nFRONTIER_VAL_FRAC    = 0.15          # held out BY SPECIES, as in CELL 6\nFRONTIER_SEED        = 0\nFRONTIER_CKPT        = \"/kaggle/working/argus_frontier.csv\"\n\n# ------------------------------------------------------------------ Part B config\nSCAN_ENABLED     = True\nSCAN_SPECIES     = PRIMARY_TARGET            # whbsho3 -- the species the project is about\nSCAN_CONTROL     = \"bkwsti\"                  # a target species with a very different niche\nSCAN_N_FILES     = 500                       # 500 x 240 s = ~33 h of unlabelled audio\nSCAN_SKIP_FIRST  = N_CALIB_BEDS + N_EVAL_BEDS   # NEVER rescan calibration/eval beds\nSCAN_TOP_K       = 60                        # clips saved for listening, per species\nSCAN_SEED        = 7\nSCAN_CKPT        = \"/kaggle/working/argus_wildscan.csv\"\nSCAN_CLIPS       = \"/kaggle/working/argus_wildscan_clips.npz\"\n\nassert EMBED_BACKEND == \"perch\", \\\n    f\"CELL 7 must start from the Perch backend, found {EMBED_BACKEND!r}\"\n\nBASE_F, CAND_F = \"stage-1 only  (p)\", \"perch only    (v)\"\nENC_F = encoders[(FEATURE, ENCODER_SEEDS[0])]   # stage 1 is IDENTICAL in every arm\n\n\n# ------------------------------------------------------------------ Part A machinery\nclass WidthStudent(nn.Module):\n    \"\"\"CELL 6's DistillStudent with the bottleneck width exposed, and an optional\n    nonlinear head.\n\n    The trunk is Enc(emb=width) -- the same four-conv-block architecture stage 1 deploys,\n    just with the final block emitting `width` channels instead of the fixed 128. Nothing\n    else about the network changes, so a difference between arms is attributable to the\n    bottleneck and not to a different model having been swapped in.\n\n    `nonlinear=True` inserts one hidden layer of the same width before the projection. It\n    exists to separate two explanations that a width sweep alone cannot: \"128 dimensions\n    cannot carry Perch's structure\" versus \"a single linear map cannot carry it, whatever\n    the width\". Those imply different fixes, so they are worth telling apart.\"\"\"\n    def __init__(s, width, out_dim, nonlinear=False):\n        super().__init__()\n        s.width = width\n        s.trunk = Enc(emb=width)\n        if nonlinear:\n            s.head = nn.Sequential(nn.Linear(width, width), nn.ReLU(), nn.Linear(width, out_dim))\n            # CELL 6's make_distill_embedder() reads `model.head.out_features` to size the\n            # empty-input return. nn.Linear has that attribute; nn.Sequential does not, so\n            # the nonlinear arm would crash the first time stage 1 proposed zero candidates\n            # for a recording -- which happens routinely, and would happen hours into the\n            # session rather than at the start. Set it explicitly so both arms satisfy the\n            # same contract.\n            s.head.out_features = out_dim\n        else:\n            s.head = nn.Linear(width, out_dim)\n        s.out_dim = out_dim\n    def forward(s, x): return s.head(s.trunk(x))\n    def embed(s, x):   return s.head(s.trunk(x))\n    def n_params(s):   return sum(p.numel() for p in s.parameters())\n\n\ndef train_width_student(specs, embs, species, width, nonlinear=False,\n                        epochs=FRONTIER_EPOCHS, batch=FRONTIER_BATCH, lr=FRONTIER_LR,\n                        val_frac=FRONTIER_VAL_FRAC, seed=FRONTIER_SEED):\n    \"\"\"Cosine-loss distillation at a given bottleneck width, validated on HELD-OUT SPECIES.\n\n    Identical protocol to CELL 6's train_distill (same loss, same species-level split, same\n    best-epoch selection on held-out species) so the width arms are comparable to CELL 6's\n    published 128-d linear result as well as to each other. The split is regenerated from\n    the same seed for every arm, so all arms see the same train/held-out species partition\n    -- otherwise a lucky partition would masquerade as a width effect.\"\"\"\n    rng = np.random.default_rng(seed)\n    uniq = np.unique(species); rng.shuffle(uniq)\n    n_val = max(1, int(round(len(uniq) * val_frac)))\n    val_sp = set(uniq[:n_val].tolist())\n    vm = np.array([s in val_sp for s in species])\n\n    Xtr = torch.tensor(specs[~vm]); Ytr = torch.tensor(embs[~vm])\n    Xva = torch.tensor(specs[vm]).to(device); Yva = torch.tensor(embs[vm]).to(device)\n\n    m = WidthStudent(width, embs.shape[1], nonlinear=nonlinear).to(device)\n    opt = optim.Adam(m.parameters(), lr)\n    sched = optim.lr_scheduler.CosineAnnealingLR(opt, T_max=epochs)\n    n = len(Xtr); best_val = -1.0; best_state = None\n    for ep in range(epochs):\n        m.train(); perm = torch.randperm(n); tot = 0.0\n        for i in range(0, n, batch):\n            idx = perm[i:i + batch]\n            xb = Xtr[idx].to(device); yb = Ytr[idx].to(device)\n            loss = (1 - nn.functional.cosine_similarity(m(xb), yb, dim=1)).mean()\n            opt.zero_grad(); loss.backward(); opt.step()\n            tot += float(loss.detach()) * len(idx)\n        sched.step()\n        m.eval()\n        with torch.no_grad():\n            vsim = float(nn.functional.cosine_similarity(m(Xva), Yva, dim=1).mean())\n        if vsim > best_val:\n            best_val = vsim\n            best_state = {k: v.detach().clone() for k, v in m.state_dict().items()}\n        if (ep + 1) % 30 == 0 or ep == 0:\n            print(f\"      ep {ep+1:3d}/{epochs}  train cos {1-tot/n:.4f}  held-out cos {vsim:.4f}\")\n    m.load_state_dict(best_state); m.eval()\n    return m, best_val\n\n\ndef _eval_backend(label, rows, extra=None):\n    \"\"\"Run the standard specificity test for whatever EMBED_BACKEND is currently active.\n\n    _TARGET_CACHE is cleared first, and this is load-bearing rather than defensive: the\n    cache key includes EMBED_BACKEND but NOT which student is bound to it, so running a\n    64-d arm and then a 256-d arm would otherwise hand the second arm the FIRST arm's\n    cached verifier matrix and report it as the second's. That failure is silent and\n    produces entirely plausible numbers, which is what makes it worth an explicit line.\"\"\"\n    _TARGET_CACHE.clear()\n    for sp in FRONTIER_EVAL_SPECIES:\n        try:\n            ctx = build_target(sp, verbose=False)\n        except Exception as ex:\n            print(f\"    skip {sp} ({label}): {type(ex).__name__}: {ex}\"); continue\n        per = []\n        for sd in FRONTIER_EVAL_SEEDS:\n            ec = eval_pass(ENC_F, ctx, sd)\n            sc = spec_pass(ENC_F, ctx, sd)\n            nc = neg_pass(ENC_F, ctx, sd)\n            m = score_eval(FUSIONS[CAND_F], ec)\n            _, _, _, _, auc = score_spec(FUSIONS[CAND_F], m[\"thr\"], sc, ctx[\"DECOYS\"])\n            per.append(dict(auc=auc, f1=m[\"f1\"], prec=m[\"p\"], rec=m[\"r\"],\n                            fa_per_hr=score_neg(FUSIONS[CAND_F], m[\"thr\"], nc)))\n        mean = {k: float(np.mean([d[k] for d in per])) for k in per[0]}\n        sd_ = {k: float(np.std([d[k] for d in per], ddof=1)) if len(per) > 1 else 0.0\n               for k in per[0]}\n        row = dict(arm=label, species=sp, n_seeds=len(FRONTIER_EVAL_SEEDS),\n                   auc=mean[\"auc\"], auc_sd=sd_[\"auc\"], f1=mean[\"f1\"], prec=mean[\"prec\"],\n                   rec=mean[\"rec\"], fa_per_hr=mean[\"fa_per_hr\"])\n        row.update(extra or {})\n        rows.append(row)\n        pd.DataFrame(rows).to_csv(FRONTIER_CKPT, index=False)\n        print(f\"    [{label:16s} {sp:9s}] AUC {mean['auc']:.3f}  F1 {mean['f1']:.3f}  \"\n              f\"FA/hr {mean['fa_per_hr']:.1f}\")\n        gc.collect()\n        try: torch.cuda.empty_cache()\n        except Exception: pass\n\n\n# ------------------------------------------------------------------ run\nfrontier_rows = []\nif SKIP_FRONTIER:\n    print(\"SKIP_FRONTIER=True -- skipping CELL 7 entirely.\")\nelse:\n    print(f\"\\n{'='*90}\\nCELL 7  PART A -- BOTTLENECK FRONTIER\\n{'='*90}\")\n    t_all = time.time()\n\n    # ---- teacher set, built ONCE and shared by every arm ----\n    t0 = time.time()\n    specs, embs, species = build_distill_set(\n        TARGET_SET, n_species=FRONTIER_SPECIES_N, recs_per_sp=FRONTIER_RECS_PER_SP,\n        crops_per_rec=FRONTIER_CROPS_PER_REC, seed=FRONTIER_SEED)\n    print(f\"[teacher set: {len(specs)} pairs, {len(np.unique(species))} species, \"\n          f\"{(time.time()-t0)/60:.1f} min]\")\n\n    # ---- Perch control, same session, same species, same seeds ----\n    if RUN_PERCH_CONTROL:\n        print(f\"\\n{'-'*90}\\nPerch control (EMBED_BACKEND='perch')\\n{'-'*90}\")\n        t0 = time.time()\n        _eval_backend(\"perch\", frontier_rows,\n                      extra=dict(width=np.nan, nonlinear=False, params=np.nan, val_cosine=np.nan))\n        print(f\"[perch control took {(time.time()-t0)/60:.1f} min -- every distilled arm \"\n              f\"below costs about the same]\")\n\n    # ---- one arm per bottleneck width ----\n    arms = [(w, False) for w in FRONTIER_WIDTHS]\n    if FRONTIER_NONLINEAR:\n        arms.append((max(FRONTIER_WIDTHS), True))\n\n    for width, nonlin in arms:\n        tag = f\"distil-{width}{'-mlp' if nonlin else ''}\"\n        print(f\"\\n{'-'*90}\\n{tag}  (bottleneck {width}-d\"\n              f\"{', nonlinear head' if nonlin else ', linear head'})\\n{'-'*90}\")\n        t0 = time.time()\n        student, val_cos = train_width_student(specs, embs, species, width, nonlinear=nonlin)\n        npar = student.n_params()\n        print(f\"    held-out-species cosine {val_cos:.4f}   params {npar:,}   \"\n              f\"{12_000_000/npar:.1f}x smaller than Perch   [{(time.time()-t0)/60:.1f} min]\")\n\n        globals()[\"DISTILL_EMBED_CTX\"] = make_distill_embedder(student)\n        try:\n            globals()[\"EMBED_BACKEND\"] = \"distilled\"\n            _eval_backend(tag, frontier_rows,\n                          extra=dict(width=width, nonlinear=nonlin, params=npar,\n                                     val_cosine=val_cos))\n        finally:\n            globals()[\"EMBED_BACKEND\"] = \"perch\"\n            globals()[\"DISTILL_EMBED_CTX\"] = None\n            _TARGET_CACHE.clear()\n        print(f\"  [{tag} total {(time.time()-t0)/60:.1f} min, \"\n              f\"{(time.time()-t_all)/60:.1f} min elapsed in CELL 7]\")\n\n    # ---- frontier report ----\n    fr = pd.DataFrame(frontier_rows)\n    if len(fr):\n        print(f\"\\n{'='*90}\\nFRONTIER\\n{'='*90}\")\n        g = fr.groupby(\"arm\", sort=False).agg(\n            auc=(\"auc\", \"mean\"), fa=(\"fa_per_hr\", \"mean\"),\n            params=(\"params\", \"first\"), cos=(\"val_cosine\", \"first\"))\n        print(f\"{'arm':<18}{'species AUC':>12}{'FA/hr':>9}{'params':>12}{'cos(teacher)':>14}\")\n        for arm, r in g.iterrows():\n            p = \"--\" if np.isnan(r[\"params\"]) else f\"{int(r['params']):,}\"\n            c = \"--\" if np.isnan(r[\"cos\"]) else f\"{r['cos']:.4f}\"\n            print(f\"{arm:<18}{r['auc']:>12.4f}{r['fa']:>9.1f}{p:>12}{c:>14}\")\n        d = g[g.index != \"perch\"]\n        if \"perch\" in g.index and len(d):\n            base = float(g.loc[\"perch\", \"auc\"])\n            best = d.auc.idxmax()\n            print(f\"\\nPerch {base:.4f}  |  best distilled '{best}' {d.auc.max():.4f}  \"\n                  f\"(gap {d.auc.max()-base:+.4f})\")\n            lin = d[~d.index.str.endswith('-mlp')]\n            if len(lin) >= 2:\n                rise = lin.auc.iloc[-1] - lin.auc.iloc[0]\n                crise = lin.cos.iloc[-1] - lin.cos.iloc[0]\n                print(f\"width {int(lin.params.notna().sum())} linear arms: \"\n                      f\"AUC {lin.auc.iloc[0]:.4f} -> {lin.auc.iloc[-1]:.4f} ({rise:+.4f}), \"\n                      f\"teacher cosine {lin.cos.iloc[0]:.4f} -> {lin.cos.iloc[-1]:.4f} ({crise:+.4f})\")\n                print(\"Read this against the pre-registered outcomes at the top of this cell:\")\n                print(\"  both rise      -> the bottleneck WAS the constraint; this is the frontier curve.\")\n                print(\"  both flat      -> it was not; CELL 6's failure is data or optimisation, not width.\")\n                print(\"  cosine rises,  -> the student reproduces Perch's geometry but loses exactly the\")\n                print(\"  AUC does not      part that separates confusable species. Report this one loudly.\")\n        print(f\"\\nwrote {FRONTIER_CKPT}  ({len(fr)} rows)\")\n\n    # =================================================================== PART B\n    if SCAN_ENABLED:\n        print(f\"\\n{'='*90}\\nCELL 7  PART B -- WILD-AUDIO SCAN\\n{'='*90}\")\n        all_paths = sorted(glob.glob(os.path.join(SS, \"*.ogg\")))\n        scan_paths = all_paths[SCAN_SKIP_FIRST:SCAN_SKIP_FIRST + SCAN_N_FILES]\n        print(f\"{len(all_paths)} soundscapes available; scanning {len(scan_paths)}, \"\n              f\"skipping the first {SCAN_SKIP_FIRST} (calibration + evaluation beds).\")\n        assert EMBED_BACKEND == \"perch\", \"PART B must run on the real Perch pipeline\"\n\n        scan_rows, clip_store = [], {}\n        for sp in (SCAN_SPECIES, SCAN_CONTROL):\n            print(f\"\\n{'-'*90}\\nscanning for {sp}\\n{'-'*90}\")\n            _TARGET_CACHE.clear()\n            ctx = build_target(sp, verbose=False)\n            # Operating threshold: the same best-F1 point every other number in this\n            # project uses, derived the same way, from planted audio -- NOT tuned on the\n            # wild files, which have no labels to tune against.\n            thr = score_eval(FUSIONS[CAND_F], eval_pass(ENC_F, ctx, SCAN_SEED))[\"thr\"]\n            print(f\"  operating threshold (best-F1, from planted eval): {thr:.4f}\")\n\n            rng = np.random.default_rng(SCAN_SEED)\n            t0 = time.time(); hits = []; secs = 0.0\n            for i, p in enumerate(scan_paths):\n                try:\n                    y = ld(p, 240.0)\n                except Exception as ex:\n                    print(f\"    skip {os.path.basename(p)}: {type(ex).__name__}\"); continue\n                secs += len(y) / SR\n                for (s0, s1, pp, vv) in detect_pv(ENC_F, ctx, y, rng):\n                    if vv >= thr:\n                        hits.append(dict(species=sp, file=os.path.basename(p),\n                                         start=float(s0), end=float(s1),\n                                         p=float(pp), v=float(vv)))\n                if i == 0:\n                    per = time.time() - t0\n                    print(f\"  [timing] first file {per:.1f}s -> ~{per*len(scan_paths)/60:.0f} min \"\n                          f\"projected for {sp}\")\n                if (i + 1) % 50 == 0:\n                    pd.DataFrame(scan_rows + hits).to_csv(SCAN_CKPT, index=False)\n                    print(f\"    {i+1}/{len(scan_paths)} files, {len(hits)} detections, \"\n                          f\"{(time.time()-t0)/60:.1f} min\")\n            hrs = secs / 3600.0\n            print(f\"  {sp}: {len(hits)} detections over {hrs:.1f} h of audio \"\n                  f\"= {len(hits)/max(hrs,1e-9):.1f}/hour   [{(time.time()-t0)/60:.1f} min]\")\n            scan_rows += hits\n\n            # top-K clips, for a human to listen to\n            top = sorted(hits, key=lambda h: -h[\"v\"])[:SCAN_TOP_K]\n            for j, h in enumerate(top):\n                try:\n                    y = ld(os.path.join(SS, h[\"file\"]), 240.0)\n                    a = max(0, int((h[\"start\"] - 0.5) * SR))\n                    b = min(len(y), int((h[\"end\"] + 0.5) * SR))\n                    clip_store[f\"{sp}_{j:03d}_v{h['v']:.3f}_{h['file'].replace('.ogg','')}\"] = y[a:b]\n                except Exception:\n                    pass\n            gc.collect()\n            try: torch.cuda.empty_cache()\n            except Exception: pass\n\n        sc = pd.DataFrame(scan_rows)\n        if len(sc):\n            sc.to_csv(SCAN_CKPT, index=False)\n            np.savez_compressed(SCAN_CLIPS, **clip_store)\n            print(f\"\\n{'='*90}\\nWILD SCAN\\n{'='*90}\")\n            summ = sc.groupby(\"species\").agg(detections=(\"v\", \"size\"), mean_v=(\"v\", \"mean\"),\n                                             max_v=(\"v\", \"max\"), files=(\"file\", \"nunique\"))\n            print(summ.to_string())\n            a = sc[sc.species == SCAN_SPECIES]; b = sc[sc.species == SCAN_CONTROL]\n            print(f\"\\n{SCAN_SPECIES} {len(a)} detections vs control {SCAN_CONTROL} {len(b)} \"\n                  f\"on the SAME {len(scan_paths)} files.\")\n            print(\"Neither number is a confirmed detection -- there is no ground truth for these\")\n            print(\"files and nobody on this project can identify the species by ear. What this\")\n            print(\"supports is a RATE, a ranked shortlist, and a comparison. If the two rates are\")\n            print(\"close, say so: that is evidence the detector is firing on bird-shaped sound\")\n            print(\"rather than on this species, and it belongs in the paper either way.\")\n            print(f\"\\nwrote {SCAN_CKPT} ({len(sc)} rows) and {SCAN_CLIPS} \"\n                  f\"({len(clip_store)} clips, load with np.load and listen)\")\n\n    print(f\"\\n[CELL 7 total {(time.time()-t_all)/60:.1f} min]\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"5cf429fe","cell_type":"markdown","source":"## After running\n\n- **This build newly writes three files** — `argus_frontier.csv` (one row per\n  species per arm: width, parameter count, teacher cosine, species AUC, F1, FA/hr),\n  `argus_wildscan.csv` (every detection above threshold, with file, timestamps and both\n  scores), and `argus_wildscan_clips.npz` (the top 60 clips per species, as raw waveforms —\n  `np.load` it and listen). Nothing else is touched: Cells 2–6 all exit early, so every\n  existing CSV in the repo stays as it is.\n- **Download all three.** The `.npz` is the one that is easy to forget and the only one\n  that cannot be regenerated without re-running the scan.\n- **Listen to the clips before believing anything in the scan CSV.** The detection counts\n  are real; whether any of them is actually *S. albiventris* is not established by this\n  cell and cannot be. Compare the study-species rate against the control rate first — if\n  they are close, that is the finding, and it belongs in the paper as one.\n- **Download the new CSV before the session ends** (or commit the notebook). Kaggle's\n  Output preview is virtualised and its plain-text extraction silently stops short of the\n  full table — a real download, or reading the preview's underlying accessibility table\n  down to its \"No more data to show\" marker, is the only reliable capture. See lab\n  notebook Day 35 for how six rows were nearly lost this way.\n- **Re-verify before publishing.** Every aggregate quoted anywhere public should be\n  recomputed from the downloaded CSV, not copied from the console. A stale console\n  number is how the README came to disagree with its own archived data (found 29 Aug).\n- Log the headline numbers — per-species AUC before/after, the winning ladder\n  config, which Perch version was used, and Cell 5's three confound verdicts — in the\n  bound logbook **the same day**.\n- If this was a `SMOKE_TEST` run: record hours-per-species-per-seed, then come back,\n  set `SMOKE_TEST = False`, and size `SEEDS` / `N_TARGETS` to the real budget before\n  the next run.\n","metadata":{}}]}