{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","file_extension":".py"},"accelerator":"GPU"},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"6e8ee4c3","cell_type":"markdown","source":"# ARGUS — Multi-Species Few-Shot Bioacoustic Detection\n\nKaggle notebook, built from the tested pipeline (`arguswala_setup_v2.py`,\n`arguswala_v2.py`, `arguswala_sweeps.py`, `arguswala_ladder.py` — 94 local assertions\npassing as of the last edit; see `_test_v2_logic.py` / `_test_ladder.py` in the project\nrepo). One notebook cell per source file, run top to bottom.\n\n## Before running — checklist\n\n1. **Notebook settings (right sidebar):** Accelerator = **GPU**, Internet = **On**.\n2. **Add data**, attach:\n   - **BirdCLEF 2024** competition data (`birdclef-2024`)\n   - **A Perch model.** Which one depends on `PERCH_VERSION` in Cell 1 (default `\"v2\"`):\n     - `PERCH_VERSION = \"v2\"` → attach `google/bird-vocalization-classifier` →\n       variation **`perch_v2`** (GPU build) — the CPU fallback `perch_v2_cpu` also works\n       if no GPU is attached, but detection is automatic either way.\n     - `PERCH_VERSION = \"v1\"` → attach `google/bird-vocalization-classifier` →\n       variation **`bird-vocalization-classifier`**, version **8**.\n   - A version mismatch between the attached model and `PERCH_VERSION` raises\n     immediately with a clear message rather than silently scoring the wrong model.\n3. **First run: leave `SMOKE_TEST = True`** in Cell 1. This limits the run to 2 species —\n   time it, write the per-species-per-seed cost in your logbook, *then* decide the full\n   `N_TARGETS` and `SEEDS` count against your actual GPU-hour budget (Kaggle free tier is\n   ~30 h/week). Flip to `False` only after that measurement.\n4. Cells 2–5 build on Cell 1's names — run in order, don't skip cells. Cell 5 does not\n   depend on Cell 3 or Cell 4 (only Cell 1+2), so it's safe to run even if you skip those.\n\nSee `ARGUS_Roadmap_Aug_Sept_2026.md` (§1b) for why the ablation ladder (Cell 4) and a\nPerch-version decision should both happen on a *few* species before committing to the\nfull multi-species run in Cell 2 — decoys are ranked by Perch-embedding proximity, so\nchanging the verifier config after Cell 2 has started makes species runs incomparable.\n","metadata":{}},{"id":"578aec5d","cell_type":"markdown","source":"## Cell 1 — Setup\n\n`arguswala_setup_v2.py`","metadata":{}},{"id":"718b738e","cell_type":"code","source":"# ===== CELL 1 v2 / SETUP -- multi-species, shot-strategy aware =====\n# Replaces arguswala_setup.py. Run BEFORE arguswala_v2.py.\n#\n# WHAT CHANGED vs v1, and why:\n#\n#  (1) MULTI-TARGET. v1 hard-wired one TARGET. Everything per-species is now behind\n#      build_target(sp), so the same pipeline evaluates N species. This pays the debt\n#      arguswala.py:375 admits to (\"one target species is not a validated finding\").\n#\n#  (2) SPECIES CACHE COMPUTED ONCE. v1 embedded 30 candidate species per target. With\n#      10 targets that is 10x redundant Perch inference -- the dominant cost. The cache\n#      is now global; build_target() only ranks against it. ~10x saving on setup.\n#\n#  (3) SHOT-SELECTION STRATEGY is now a variable, not an assumption. v1 always took the\n#      5 highest-`rating` recordings. DCASE always takes the first 5. Nolasco et al.\n#      (2023) sec.4.1 report that unrepresentative support sets were the dominant failure\n#      cause their expert annotators identified, and call better selection an open\n#      problem. So it becomes an axis we measure instead of a default we inherit.\n#\n#  (4) PER-SPECIES COVARIATES are recorded at build time (decoy proximity, stereotypy,\n#      spectral centre/bandwidth, call duration, n recordings). Nolasco et al. tried\n#      multivariate regression to find what predicts task difficulty and FAILED to find\n#      any consistent factor. These covariates are free at runtime and impossible to\n#      reconstruct afterwards -- logging them is what makes that analysis possible later.\n#\n#  (5) The held-out pool is now derived from the shots actually chosen, not from a\n#      positional slice. Under a non-'first' strategy, files[N_SHOT:] would have leaked\n#      support recordings into the evaluation pool.\nimport os, glob, time\nimport numpy as np\nimport pandas as pd\nimport librosa\nimport tensorflow_hub as hub\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.cluster import KMeans\n\n# ---------------------------------------------------------------- ablation ladder\n# Each rung is independently switchable so marginal contributions are measurable.\n# Rung 0 is the original pipeline; nothing below changes it when all flags are False.\nVERIFIER = \"logreg\"        # 'logreg' (rung 0) | 'proto' (rung 1, stage 2)\n\n# ------------------------------------------------------------------ configuration\nSR = 22050\nSEG_SEC = 1.0\nCLEN = int(SEG_SEC * SR)\n\nN_SHOT          = 5\nSHOT_STRATEGY   = \"rated\"     # 'rated' | 'first' | 'random' | 'diverse'\nN_DECOYS        = 4\nN_CANDIDATES    = 30          # non-target species considered for decoy ranking\nCALIB_PER_SPECIES = 6\nDECOY_RECORDINGS  = 20\nN_CALIB_BEDS, N_EVAL_BEDS = 4, 10\nSNR_LEVELS = (9, 6, 3, 0, -3, -6)\n\n# Targets are PRE-REGISTERED: chosen by a rule, before any result is seen.\n# Rule: species with >= MIN_RECS recordings, ranked by recording count, top N_TARGETS,\n# with the original study species forced in first. Deterministic -- no cherry-picking.\nPRIMARY_TARGET = \"whbsho3\"    # White-bellied Sholakili, the original study species\nN_TARGETS      = 10\nMIN_RECS       = N_SHOT + 4   # 5 shots + >=4 genuinely held out\n\nSMOKE_TEST = False            # 15 Aug 2026: timing measured (10 min/species/3seeds),\n                              # config locked (PCEN/logreg/no-tfa/no-hn all confirmed\n                              # winners in the ladder), noise floor known (0.025-0.035).\n                              # Centrepiece run: full N_TARGETS species, no more reason\n                              # to stay on the 2-species smoke test.\n\n# Added 14 Aug 2026, after the first real run: whbsho3 stage-1-alone AUC came back 0.517\n# in a run that excludes 2 species from the encoder bank (this multi-target pipeline),\n# vs. the original single-species baseline of 0.459, which only ever excluded whbsho3.\n# ENCODER_SEEDS (widened separately, in CELL 2) tests training-run-to-training-run noise\n# on a FIXED bank. This flag tests the OTHER candidate cause: does the bank's COMPOSITION\n# itself change when only 1 species is excluded instead of 2? Set True to force\n# TARGET_LIST down to [PRIMARY_TARGET] alone -- the exact exclusion set the original\n# baseline used -- so build_banks() in CELL 2 produces a directly comparable bank.\n# Re-run CELL 1 then CELL 2 after flipping this; no kernel restart needed, everything\n# downstream is rebuilt from these globals. CELL 3/4 are not needed for this comparison.\nSINGLE_TARGET_COMPARISON = False\n\nroot = os.path.dirname(glob.glob('/kaggle/input/**/train_metadata.csv', recursive=True)[0])\nAUD = os.path.join(root, 'train_audio')\nSS  = os.path.join(root, 'unlabeled_soundscapes')\nmeta = pd.read_csv(os.path.join(root, 'train_metadata.csv'))\n\n# Recordist identity column, for the recordist-disjoint confound check (CELL 5, roadmap\n# \"owed\" list: rules out \"did it learn the microphone, not the bird\"). BirdCLEF inherits\n# Xeno-Canto's schema, where the recordist's name is the 'author' field -- but that is a\n# CONVENTION, not something this script has verified against the actual attached dataset,\n# so it is printed here rather than assumed silently. If build_target(recordist_disjoint=\n# True) raises a KeyError, check this list against the real column name and fix RECORDIST_COL.\nRECORDIST_COL = \"author\"\nprint(f\"train_metadata.csv columns: {meta.columns.tolist()}\")\nprint(f\"  RECORDIST_COL = {RECORDIST_COL!r} \"\n      f\"({'present' if RECORDIST_COL in meta.columns else 'NOT FOUND -- fix before using recordist_disjoint'})\")\n\n# ------------------------------------------------------------------ audio helpers\ndef rms(x):\n    return float(np.sqrt(np.mean(np.square(x, dtype=np.float64)) + 1e-12))\n\ndef ld(p, dur=None):\n    y, _ = librosa.load(p, sr=SR, mono=True, duration=dur)\n    return y.astype('float32')\n\ndef l2n(x):\n    x = np.asarray(x, dtype=\"float32\")\n    return x / (np.linalg.norm(x, axis=-1, keepdims=True) + 1e-9)\n\ndef cos(a, b):\n    return float(a @ b / ((np.linalg.norm(a) + 1e-9) * (np.linalg.norm(b) + 1e-9)))\n\ndef loudest_offset(y, n=CLEN):\n    if len(y) <= n: return 0\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    return int(np.argmax(cs[n:] - cs[:-n]))\n\ndef peak_windows(y, k=1, n=CLEN):\n    \"\"\"Top-k non-overlapping highest-energy windows of length n.\"\"\"\n    if len(y) <= n: return [np.pad(y, (0, n - len(y))).astype('float32')]\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    w = cs[n:] - cs[:-n]\n    out, taken = [], []\n    for _ in range(k):\n        pick = next((int(i) for i in np.argsort(w)[::-1] if all(abs(int(i) - t) >= n for t in taken)), None)\n        if pick is None: break\n        taken.append(pick); out.append(y[pick:pick + n].astype('float32'))\n    return out\n\ndef build_support_views(recordings, shifts=(-0.25, 0.0, 0.25)):\n    out = []\n    for y in recordings:\n        base = loudest_offset(y)\n        for shift in shifts:\n            a = int(np.clip(base + shift * SR, 0, max(0, len(y) - CLEN)))\n            w = y[a:a + CLEN]\n            if len(w) < CLEN: w = np.pad(w, (0, CLEN - len(w)))\n            out.append(w.astype('float32'))\n    return out\n\ndef place(y, pool, rng, spans, snr_db, gap_s=0.4):\n    \"\"\"Mix one random call from `pool` into `y` at a free position -> centre in seconds.\"\"\"\n    for _ in range(80):\n        pos = int(rng.integers(0, len(y) - CLEN))\n        if any(pos < e + int(gap_s*SR) and pos + CLEN > s - int(gap_s*SR) for s, e in spans):\n            continue\n        c = pool[int(rng.integers(len(pool)))][:CLEN]\n        if len(c) < CLEN: c = np.pad(c, (0, CLEN - len(c)))\n        g = rms(y[pos:pos+CLEN]) * (10 ** (float(snr_db)/20)) / (rms(c) + 1e-9)\n        y[pos:pos+CLEN] += c.astype('float32') * g\n        spans.append((pos, pos + CLEN))\n        return (pos + CLEN/2) / SR\n    return None\n\n# ------------------------------------------------------------------ Perch\n# PERCH 2.0 UPGRADE (Aug 2025 release, arXiv 2508.04665). Three things changed and each\n# one silently breaks the v1 code path, so all three are handled explicitly. Verified\n# against the official loader: google-research/perch-hoplite, zoo/taxonomy_model_tf.py.\n#\n#  (a) NO `infer_tf`. Perch 1.0 was a TF-Jax export exposing model.infer_tf(). Perch 2.0\n#      exposes only the standard serving signature. perch-hoplite branches on exactly\n#      this: `elif hasattr(self.model, 'infer_tf') ... else: # Perch v2 style\n#      outputs = self.model.signatures['serving_default'](inputs=rebatched_audio)`.\n#      Calling infer_tf on v2 raises AttributeError.\n#\n#  (b) EMBEDDING WIDTH 1280 -> 1536 (EfficientNet-B1 -> B3). Probed, never assumed.\n#      v2 also emits a `spatial_embedding` (5,3,1536) head that must be ignored --\n#      `embedding` is the pooled mean vector we want.\n#\n#  (c) AUDIO IS PEAK-NORMALISED BEFORE INFERENCE. The official wrapper removes DC offset\n#      and scales each window to a peak of 0.25 (zoo_interface.normalize_audio,\n#      target_peak=0.25). ARGUS v1 fed raw waveforms to Perch and did NEITHER -- so\n#      Perch was being given amplitudes outside the distribution it is served with.\n#      This affects the EXISTING v1 numbers too, not just the upgrade.\n#      Note it does NOT disturb the SNR sweep: peak norm is a per-window scalar multiply,\n#      and SNR is a ratio within the window, so planted-call SNR is preserved exactly.\nPERCH_VERSION   = \"v2\"     # \"v2\" | \"v1\"  -- recorded in the results CSV\nPERCH_PEAK_NORM = 0.25     # None to disable (v1 behaviour); 0.25 matches the official wrapper\nPERCH_BACKEND   = \"tf\"     # \"tf\" | \"onnx\" -- see below. Default \"tf\" -> ZERO behaviour\n                           # change to the tested GPU pipeline unless explicitly opted in.\n\nPERCH_SR, PERCH_WIN = 32000, 5.0     # unchanged between versions\nPERCH_LEN = int(PERCH_WIN * PERCH_SR)\nPERCH_BATCH = 16\n\n_V2_SLUGS = (\"perch_v2\", \"perch_v2_cpu\")\ndef _find_local_perch(prefer_v2):\n    \"\"\"Kaggle attaches models as input dirs. Pick one matching the requested version.\"\"\"\n    hits = [os.path.dirname(p) for p in glob.glob('/kaggle/input/**/saved_model.pb', recursive=True)\n            if 'bird-vocalization' in p.lower() or 'perch' in p.lower()]\n    if not hits: return None\n    is_v2 = lambda p: any(s in p.lower() for s in _V2_SLUGS)\n    pool = [p for p in hits if is_v2(p) == prefer_v2] or hits\n    def _ver(p):\n        try: return int(os.path.basename(p))\n        except ValueError: return -1\n    return max(pool, key=_ver)\n\ndef _remote_handle():\n    if PERCH_VERSION != \"v2\":\n        # perch_8 in perch-hoplite's config; v8 is the last 1280-d release.\n        return (\"https://www.kaggle.com/models/google/bird-vocalization-classifier\"\n                \"/tensorFlow2/bird-vocalization-classifier/8\")\n    # perch-hoplite ships separate GPU and CPU builds at different version numbers.\n    try:\n        import tensorflow as tf\n        gpu = bool(tf.config.list_physical_devices('GPU'))\n    except Exception:\n        gpu = False\n    slug, ver = (\"perch_v2\", 2) if gpu else (\"perch_v2_cpu\", 1)\n    return (\"https://www.kaggle.com/models/google/bird-vocalization-classifier\"\n            f\"/tensorFlow2/{slug}/{ver}\")\n\n# ---------------------------------------------------------------- ONNX backend (opt-in)\n# For CPU-constrained / edge-hardware experiments (roadmap: ARGUS_Hardware_Deployment_\n# Roadmap.md), not for the main GPU pipeline. Uses a pre-converted, community-published,\n# Apache-2.0 ONNX export (justinchuby/Perch-onnx on Hugging Face) rather than converting\n# ourselves -- same I/O signature, independently verified against the real 413MB weights\n# on 13 Aug 2026: input 'inputs' [batch,160000], output 'embedding' [batch,1536], real\n# inference measured at ~156ms/clip (batch=16, 4 CPU threads) on a laptop CPU -- NOT the\n# target embedded device, a reference point only. Peak-norm was confirmed to materially\n# change the real embedding (mean abs delta 0.065), i.e. the v1 peak-norm bug was real,\n# not theoretical, on the actual weights.\n_ONNX_HF_URL = \"https://huggingface.co/justinchuby/Perch-onnx/resolve/main/perch_v2_no_dft.onnx\"\n\ndef _find_local_onnx():\n    hits = (glob.glob('/kaggle/input/**/perch_v2_no_dft*.onnx', recursive=True)\n            or glob.glob('/kaggle/input/**/perch_v2*.onnx', recursive=True))\n    return hits[0] if hits else None\n\ndef _load_onnx_perch():\n    import onnxruntime as ort\n    path = _find_local_onnx()\n    if path is None:\n        # Falls back to a live download (~413MB) if nothing is attached as a Kaggle\n        # input -- slow and wasteful on repeated runs. Attaching a Kaggle dataset\n        # mirror (e.g. tuckerarrants/perch-v2-no-dft-onnx) is the fast path.\n        import requests\n        path = \"/kaggle/working/perch_v2_no_dft.onnx\"\n        print(f\"  no local ONNX Perch found; downloading from {_ONNX_HF_URL} ...\")\n        r = requests.get(_ONNX_HF_URL, stream=True, timeout=300)\n        with open(path, \"wb\") as f:\n            for chunk in r.iter_content(chunk_size=1 << 20):\n                f.write(chunk)\n    so = ort.SessionOptions(); so.intra_op_num_threads = 4\n    sess = ort.InferenceSession(path, sess_options=so, providers=[\"CPUExecutionProvider\"])\n    in_name = sess.get_inputs()[0].name\n    out_names = [o.name for o in sess.get_outputs()]\n    print(f\"  ONNX Perch loaded from {path}  |  input='{in_name}'  outputs={out_names}\")\n    return sess, in_name, out_names\n\nif PERCH_BACKEND == \"onnx\":\n    _ort_sess, _ort_in, _ort_out_names = _load_onnx_perch()\n    _HAS_INFER_TF = False; _SERVING = None; perch = None   # not used on this path\n    print(f\"  calling convention: onnxruntime (CPU)  |  peak-norm: {PERCH_PEAK_NORM}\")\nelse:\n    PERCH_HANDLE = _find_local_perch(PERCH_VERSION == \"v2\") or _remote_handle()\n    print(f\"Perch {PERCH_VERSION} source: {PERCH_HANDLE}\")\n    perch = hub.load(PERCH_HANDLE)\n\n    # Which calling convention does THIS artefact expose? Detected, not assumed -- so a\n    # mis-attached Kaggle input degrades to a clear message instead of a confusing crash.\n    _HAS_INFER_TF = hasattr(perch, \"infer_tf\")\n    _SERVING = None\n    if not _HAS_INFER_TF:\n        try:\n            _SERVING = perch.signatures[\"serving_default\"]\n        except Exception as ex:\n            raise RuntimeError(\n                \"Loaded Perch artefact exposes neither infer_tf nor a 'serving_default' \"\n                f\"signature ({type(ex).__name__}). Check the attached Kaggle model.\"\n            ) from ex\n    print(f\"  calling convention: {'infer_tf (Perch 1.x)' if _HAS_INFER_TF else 'serving_default (Perch 2.x)'}\"\n          f\" | peak-norm: {PERCH_PEAK_NORM}\")\n_perch_batched = True\n\ndef _probe_embed_dim():\n    try:\n        return int(_perch_run(np.zeros((1, PERCH_LEN), dtype=\"float32\")).shape[-1])\n    except Exception as ex:\n        print(\"  could not probe embedding dim, assuming 1280:\", type(ex).__name__)\n        return 1280\n\ndef to32k(y):\n    return librosa.resample(np.asarray(y, dtype=\"float32\"), orig_sr=SR, target_sr=PERCH_SR)\n\ndef _peak_norm(batch):\n    \"\"\"DC-remove and scale each window to PERCH_PEAK_NORM -- zoo_interface.normalize_audio.\"\"\"\n    if PERCH_PEAK_NORM is None: return batch\n    b = np.asarray(batch, dtype=\"float32\").copy()\n    b -= np.mean(b, axis=-1, keepdims=True)\n    pk = np.max(np.abs(b), axis=-1, keepdims=True)\n    b = np.divide(b, pk, where=(pk > 0.0))\n    return (b * PERCH_PEAK_NORM).astype(\"float32\")\n\ndef _extract_embedding(out):\n    \"\"\"Perch 2.x -> dict with 'embedding' (pooled) plus 'spatial_embedding'/'frontend'/\n    logit heads. Perch 1.x -> dict with 'embedding', or a bare (logits, embeddings) tuple.\"\"\"\n    if hasattr(out, \"keys\"):\n        if \"embedding\" not in out:\n            raise KeyError(f\"no 'embedding' head in Perch output; got {sorted(out.keys())}\")\n        emb = out[\"embedding\"]\n    else:\n        emb = out[1]\n    return np.asarray(emb.numpy() if hasattr(emb, \"numpy\") else emb)\n\ndef _infer(batch):\n    batch = _peak_norm(batch)\n    if PERCH_BACKEND == \"onnx\":\n        result = _ort_sess.run(None, {_ort_in: batch})\n        out = dict(zip(_ort_out_names, result))   # verified real keys: embedding/\n                                                   # spatial_embedding/spectrogram/label\n    else:\n        out = perch.infer_tf(batch) if _HAS_INFER_TF else _SERVING(inputs=batch)\n    return _extract_embedding(out)\n\ndef _perch_run(waves):\n    global _perch_batched\n    n = len(waves)\n    if n == 0: return np.zeros((0, EMBED_DIM if \"EMBED_DIM\" in globals() else 1280), dtype=\"float32\")\n    if _perch_batched:\n        try:\n            out = []\n            for i in range(0, n, PERCH_BATCH):\n                chunk = waves[i:i+PERCH_BATCH]; k = len(chunk)\n                if k < PERCH_BATCH:\n                    chunk = np.concatenate([chunk, np.zeros((PERCH_BATCH-k, PERCH_LEN), dtype=\"float32\")])\n                out.append(_infer(chunk)[:k])\n            return np.concatenate(out)\n        except Exception as ex:\n            print(\"batched Perch inference failed, falling back to per-clip:\", type(ex).__name__)\n            _perch_batched = False\n    return np.stack([_infer(w[np.newaxis, :])[0] for w in waves])\n\ndef perch_embed(waveforms):\n    \"\"\"Isolated clips -> (n, D). CENTRED zero-padding, Perch's documented convention.\"\"\"\n    ws = []\n    for w in waveforms:\n        w = to32k(w)\n        if len(w) >= PERCH_LEN:\n            w = w[:PERCH_LEN]\n        else:\n            pad = PERCH_LEN - len(w); lo = pad // 2\n            w = np.pad(w, (lo, pad - lo))\n        ws.append(w.astype(\"float32\"))\n    return _perch_run(np.stack(ws)) if ws else np.zeros((0, EMBED_DIM), dtype=\"float32\")\n\ndef perch_embed_ctx(y32, centres_sec):\n    \"\"\"Real 5s windows cut from a 32kHz recording, centred on each event time.\"\"\"\n    ws = []\n    for c in centres_sec:\n        a = int(round(c * PERCH_SR)) - PERCH_LEN // 2\n        a = max(0, min(a, max(0, len(y32) - PERCH_LEN)))\n        w = y32[a:a + PERCH_LEN]\n        if len(w) < PERCH_LEN:\n            pad = PERCH_LEN - len(w); lo = pad // 2\n            w = np.pad(w, (lo, pad - lo))\n        ws.append(w.astype(\"float32\"))\n    return _perch_run(np.stack(ws)) if ws else np.zeros((0, EMBED_DIM), dtype=\"float32\")\n\nEMBED_DIM = _probe_embed_dim()\n_DETECTED = {1280: \"v1\", 1536: \"v2\"}.get(EMBED_DIM, \"unknown\")\nprint(f\"Perch embedding dim: {EMBED_DIM}  (detected {_DETECTED}, requested {PERCH_VERSION})\")\nif _DETECTED != PERCH_VERSION:\n    # Wrong artefact attached is the likeliest cause, and it would otherwise produce\n    # perfectly plausible numbers from the wrong model. Refuse rather than mislead.\n    raise RuntimeError(\n        f\"Perch version mismatch: requested {PERCH_VERSION} but the loaded model emits \"\n        f\"{EMBED_DIM}-d embeddings ({_DETECTED}). Loaded from: {PERCH_HANDLE}\\n\"\n        f\"  -> On Kaggle, attach the model variation matching PERCH_VERSION:\\n\"\n        f\"     v2: bird-vocalization-classifier/tensorFlow2/perch_v2 (GPU, v2) \"\n        f\"or perch_v2_cpu (CPU, v1)\\n\"\n        f\"     v1: bird-vocalization-classifier/tensorFlow2/bird-vocalization-classifier (v8)\")\n\n# ------------------------------------------------- pre-registered target selection\n_counts = meta.primary_label.value_counts()\n_eligible = [sp for sp in _counts.index if _counts[sp] >= MIN_RECS]\nif PRIMARY_TARGET not in _eligible:\n    raise ValueError(f\"{PRIMARY_TARGET} has only {_counts.get(PRIMARY_TARGET, 0)} recordings \"\n                     f\"(need >= {MIN_RECS})\")\nif SINGLE_TARGET_COMPARISON:\n    TARGET_LIST = [PRIMARY_TARGET]\nelse:\n    TARGET_LIST = [PRIMARY_TARGET] + [sp for sp in _eligible if sp != PRIMARY_TARGET][:N_TARGETS - 1]\n    if SMOKE_TEST:\n        TARGET_LIST = TARGET_LIST[:2]\nTARGET_SET = set(TARGET_LIST)     # <-- consumed by CELL 2 to exclude ALL targets from the\n                                  #     encoder bank. v1 hard-coded \"whbsho3\" there, which\n                                  #     silently trained the encoder on every OTHER target.\nprint(f\"\\nPRE-REGISTERED TARGETS (n={len(TARGET_LIST)}, smoke_test={SMOKE_TEST}, \"\n      f\"single_target_comparison={SINGLE_TARGET_COMPARISON}) \"\n      f\"-- copy this list into the logbook BEFORE running:\")\nfor sp in TARGET_LIST:\n    print(f\"   {sp}  ({_counts[sp]} recordings)\")\n\n# ------------------------------------------------- global species cache (ONCE)\n_cand = (meta[~meta.primary_label.isin(TARGET_SET)]['primary_label']\n         .value_counts().head(N_CANDIDATES).index.tolist())\nSPECIES_CACHE = {}\n_t0 = time.time()\nfor sp in _cand:\n    calls = []\n    for f in meta[meta.primary_label == sp]['filename'].tolist()[:CALIB_PER_SPECIES]:\n        try: calls += peak_windows(ld(os.path.join(AUD, f), 30), k=1)\n        except Exception: pass\n    if not calls: continue\n    SPECIES_CACHE[sp] = {\"calls\": calls, \"emb\": l2n(perch_embed(calls)).mean(0)}\nprint(f\"species cache: {len(SPECIES_CACHE)}/{len(_cand)} candidates embedded \"\n      f\"in {time.time()-_t0:.0f}s (computed once, reused by every target)\")\n\n_decoy_pool_cache = {}\ndef _deep_decoy_pool(sp):\n    if sp not in _decoy_pool_cache:\n        pool = []\n        for f in meta[meta.primary_label == sp]['filename'].tolist()[:DECOY_RECORDINGS]:\n            try: pool += peak_windows(ld(os.path.join(AUD, f), 30), k=1)\n            except Exception: pass\n        _decoy_pool_cache[sp] = pool or SPECIES_CACHE[sp][\"calls\"]\n    return _decoy_pool_cache[sp]\n\n# ------------------------------------------------------------------ soundscape beds\nbed_paths = sorted(glob.glob(os.path.join(SS, \"*.ogg\")))[:N_CALIB_BEDS + N_EVAL_BEDS]\nall_beds = [ld(p, 240.0) for p in bed_paths]\ncalib_beds, eval_beds = all_beds[:N_CALIB_BEDS], all_beds[N_CALIB_BEDS:]\nBED_SEC = 240.0\nprint(f\"beds: {len(calib_beds)} calibration + {len(eval_beds)} evaluation (disjoint)\")\n\n# ------------------------------------------------------------ shot selection\ndef choose_shots(files, ratings, embs, n_shot, strategy, rng):\n    \"\"\"-> indices into `files`. Each strategy is a different answer to the open question\n    Nolasco et al. sec.4.1 raise: which 5 examples should a deployed system be given?\n      rated   : highest Xeno-Canto rating (ARGUS v1 behaviour -- best-quality audio)\n      first   : first n in metadata order (DCASE Task 5 behaviour -- no curation)\n      random  : uniform sample (control)\n      diverse : greedy max-min distance in Perch space (maximum coverage of the\n                call repertoire -- the thing the expert annotators said was missing)\n    \"\"\"\n    n = len(files)\n    if strategy == \"rated\":\n        return list(np.argsort(-np.asarray(ratings, dtype=float))[:n_shot])\n    if strategy == \"first\":\n        return list(range(min(n_shot, n)))\n    if strategy == \"random\":\n        return list(rng.permutation(n)[:n_shot])\n    if strategy == \"diverse\":\n        E = l2n(embs)\n        picked = [int(np.argmax(np.asarray(ratings, dtype=float)))]   # deterministic seed point\n        while len(picked) < min(n_shot, n):\n            d = np.min(1.0 - E @ E[picked].T, axis=1)   # cosine distance to nearest picked\n            d[picked] = -np.inf\n            picked.append(int(np.argmax(d)))\n        return picked\n    raise ValueError(f\"unknown shot strategy: {strategy}\")\n\n# ------------------------------------------------------------ stage-2 verifiers\nclass PrototypicalProbe:\n    \"\"\"Prototype-distance classifier. Drop-in for LogisticRegression (fit/predict_proba).\n\n    WHY: Perch 2.0's own benchmark (Table 3, BEANS) reports prototypical probing beating\n    linear probing on DETECTION tasks by +0.078 mAP (0.426 -> 0.504). ARGUS stage 2 is a\n    detection task, and currently uses a linear probe.\n\n    K prototypes per class rather than one mean, following the ProtoPNet lineage Perch 2.0\n    adopts (4 prototypes/class, prediction = max activation across them). This is also the\n    direct answer to Nolasco et al. sec.4.1: a single averaged prototype \"implies that the\n    examples are in some sense all of one kind\" -- false for a species with several call\n    types, which is exactly ARGUS's failure mode.\n\n    Scores are cosine similarities (inputs are L2-normalised upstream), turned into\n    calibrated probabilities by a fitted temperature. Calibration matters here because the\n    fusion rules multiply p*v -- an uncalibrated v that saturates to 0/1 becomes a hard\n    veto and destroys ranking, which is the same trap that forced C=0.1 on the LogReg.\n    \"\"\"\n\n    def __init__(self, n_prototypes=4, random_state=0, class_weight=\"balanced\"):\n        self.n_prototypes = n_prototypes\n        self.random_state = random_state\n        self.class_weight = class_weight\n\n    def _protos(self, Xc):\n        k = int(min(self.n_prototypes, len(Xc)))\n        if k <= 1:\n            return l2n(Xc.mean(0, keepdims=True))\n        km = KMeans(n_clusters=k, n_init=10, random_state=self.random_state).fit(Xc)\n        return l2n(km.cluster_centers_)\n\n    def _scores(self, X):\n        \"\"\"-> (n, 2): max cosine similarity to each class's prototype set.\"\"\"\n        Xn = l2n(X)\n        return np.stack([(Xn @ self.protos_[c].T).max(1) for c in (0, 1)], axis=1)\n\n    @staticmethod\n    def _softmax(z):\n        z = z - z.max(axis=1, keepdims=True)\n        e = np.exp(z)\n        return e / e.sum(axis=1, keepdims=True)\n\n    def _fit_temperature(self, S, y, w, eps=0.05):\n        \"\"\"Temperature by weighted log-loss with LABEL SMOOTHING.\n\n        Smoothing is not cosmetic here. The calibration set separates almost perfectly in\n        Perch space, so an unsmoothed objective drives T -> 0 and saturates predict_proba\n        to hard 0/1. That turns v into a veto rather than a score, which destroys the\n        ranking that AP and the p*v fusion rules depend on -- the identical failure mode\n        that forced C=0.1 on the LogisticRegression verifier. With targets smoothed to\n        1-eps the optimum is finite and the probabilities stay soft.\n        \"\"\"\n        best_T, best_ll = 1.0, np.inf\n        for T in np.geomspace(0.02, 3.0, 60):\n            p_true = np.clip(self._softmax(S / T)[np.arange(len(y)), y], 1e-9, 1.0)\n            p_false = np.clip(1.0 - p_true, 1e-9, 1.0)\n            ll = float(-(w * ((1.0 - eps) * np.log(p_true)\n                              + eps * np.log(p_false))).sum() / w.sum())\n            if ll < best_ll:\n                best_T, best_ll = float(T), ll\n        return best_T\n\n    def fit(self, X, y):\n        X = np.asarray(X, dtype=\"float32\"); y = np.asarray(y).astype(int)\n        self.protos_ = {c: self._protos(X[y == c]) for c in (0, 1)}\n        if self.class_weight == \"balanced\":\n            cw = {c: len(y) / (2.0 * max(1, (y == c).sum())) for c in (0, 1)}\n        else:\n            cw = {0: 1.0, 1: 1.0}\n        w = np.array([cw[int(v)] for v in y], dtype=\"float64\")\n        self.temperature_ = self._fit_temperature(self._scores(X), y, w)\n        return self\n\n    def predict_proba(self, X):\n        return self._softmax(self._scores(X) / self.temperature_)\n\n    def score(self, X, y):\n        return float((self.predict_proba(X).argmax(1) == np.asarray(y).astype(int)).mean())\n\n\ndef _make_verifier():\n    if VERIFIER == \"logreg\":\n        # C deliberately small: 1280/1536-d embeddings with ~100 samples separate perfectly\n        # at C=1, saturating predict_proba to 0/1 and turning v into a hard veto.\n        return LogisticRegression(C=0.1, max_iter=2000, class_weight=\"balanced\")\n    if VERIFIER == \"proto\":\n        return PrototypicalProbe(n_prototypes=4, random_state=0)\n    raise ValueError(f\"unknown VERIFIER: {VERIFIER}\")\n\n\n# ------------------------------------------------------------ per-species covariates\ndef _covariates(support_views, own_embs, decoy_sims):\n    \"\"\"Recorded so the difficulty-factor regression Nolasco et al. could not resolve\n    becomes possible later. Free now; unreconstructable after the fact.\"\"\"\n    E = l2n(own_embs)\n    if len(E) > 1:\n        S = E @ E.T\n        iu = np.triu_indices(len(E), k=1)\n        stereotypy = float(S[iu].mean())        # high = calls resemble each other\n    else:\n        stereotypy = float('nan')\n    cents, bws, durs = [], [], []\n    for w in support_views:\n        cents.append(float(np.mean(librosa.feature.spectral_centroid(y=w, sr=SR))))\n        bws.append(float(np.mean(librosa.feature.spectral_bandwidth(y=w, sr=SR))))\n        e = np.square(w, dtype=np.float64)\n        thr = 0.1 * e.max() if e.max() > 0 else 0.0\n        durs.append(float((e > thr).sum()) / SR)   # seconds above 10% of peak energy\n    return dict(\n        stereotypy=stereotypy,\n        spectral_centroid_hz=float(np.mean(cents)),\n        spectral_bandwidth_hz=float(np.mean(bws)),\n        call_duration_s=float(np.mean(durs)),\n        decoy_sim_top1=float(max(decoy_sims)) if decoy_sims else float('nan'),\n        decoy_sim_mean=float(np.mean(decoy_sims)) if decoy_sims else float('nan'),\n    )\n\n# ------------------------------------------------------------ per-target builder\n# Everything except the verifier fit is deterministic given (species, n_shot, strategy,\n# seed) -- and it is nearly all Perch inference, the dominant cost. The ablation ladder\n# rebuilds targets once per rung, so without this cache every rung would re-embed the same\n# audio. Cached across rungs; only the (cheap) classifier is refitted.\n_TARGET_CACHE = {}\n\ndef build_target(sp, n_shot=N_SHOT, strategy=SHOT_STRATEGY, seed=0, verbose=True,\n                  decoy_mode=\"hard\", recordist_disjoint=False):\n    \"\"\"Everything v1 did at module level for one hard-coded species, now per species.\n    Returns a self-contained context dict consumed by CELL 2.\n\n    decoy_mode selects WHICH species (out of every candidate ranked by Perch-embedding\n    similarity to this target) become the N_DECOYS used at test time -- roadmap 3.3:\n      'hard'   (default) -- the top N_DECOYS nearest. Unchanged behaviour; every result\n                published before this flag existed used this mode.\n      'medium' -- ranks N_DECOYS..2*N_DECOYS-1 (5th-8th nearest at N_DECOYS=4).\n      'random' -- N_DECOYS drawn uniformly from the full non-target candidate pool,\n                  seeded from this call's own `seed` so it's reproducible, not a\n                  fresh draw on every call.\n    Only what a run is SCORED against changes; the verifier is still trained against\n    every non-decoy candidate as negatives regardless of mode (calib_neg_species below),\n    so 'easier' decoys are a harder TEST, not a easier training signal.\n\n    recordist_disjoint -- confound-closing run 3 (roadmap \"owed\" list): \"did it learn the\n    microphone, not the bird?\" Default False reproduces every previously-published number\n    byte-for-byte (held-out pool = every recording not used as a shot, regardless of who\n    recorded it). True additionally excludes, from the held-out pool, any recording whose\n    RECORDIST_COL matches one of the n_shot support recordings -- so if a species'\n    apparent AUC were actually partly \"recognise this recordist's mic/room\", a\n    recordist-disjoint held-out pool would show it as a drop. Can raise ValueError if a\n    species has too few distinct recordists (the held-out pool empties out) -- that itself\n    is worth logging, not a bug to route around.\"\"\"\n    ck = (sp, int(n_shot), strategy, int(seed), decoy_mode, bool(recordist_disjoint))\n    if ck in _TARGET_CACHE:\n        c = _TARGET_CACHE[ck]\n        clf = _make_verifier().fit(c[\"X\"], c[\"yv\"])\n        cov = dict(c[\"cov\"]); cov[\"clf_train_acc\"] = float(clf.score(c[\"X\"], c[\"yv\"]))\n        cov[\"verifier\"] = VERIFIER\n        if verbose:\n            print(f\"  [{sp}] cached build reused; verifier={VERIFIER} \"\n                  f\"(train acc {cov['clf_train_acc']:.2f})\")\n        return dict(species=sp, support=c[\"support\"], target_pool=c[\"target_pool\"],\n                    PROTO=c[\"PROTO\"], DECOYS=c[\"DECOYS\"], decoy_pools=c[\"decoy_pools\"],\n                    perch_clf=clf, covariates=cov, neg_calls=c[\"neg_calls\"])\n    rng = np.random.default_rng(seed)\n    rows = meta[meta.primary_label == sp]\n    files   = rows['filename'].tolist()\n    ratings = rows['rating'].tolist() if 'rating' in rows else [0.0] * len(files)\n    if recordist_disjoint and RECORDIST_COL not in rows:\n        raise KeyError(f\"recordist_disjoint=True but {RECORDIST_COL!r} is not a column \"\n                        f\"of train_metadata.csv; have {rows.columns.tolist()}\")\n    authors = (rows[RECORDIST_COL].tolist() if recordist_disjoint\n               else [None] * len(files))\n    if len(files) < n_shot + 2:\n        raise ValueError(f\"not enough recordings for {sp}: {len(files)}\")\n\n    raw = [ld(os.path.join(AUD, f), 30) for f in files]\n    # 'diverse' needs one embedding per recording to pick a spread; others don't.\n    rec_embs = (perch_embed([peak_windows(y, k=1)[0] for y in raw])\n                if strategy == \"diverse\" else np.zeros((len(raw), 1)))\n\n    shot_idx = choose_shots(files, ratings, rec_embs, n_shot, strategy, rng)\n    shot_set = set(int(i) for i in shot_idx)\n\n    support = build_support_views([raw[i] for i in shot_idx])\n    # Held-out pool = every recording NOT used as a shot. v1 sliced files[N_SHOT:],\n    # which under any non-'first' strategy would have leaked support into evaluation.\n    excluded_idx = set(shot_set)\n    if recordist_disjoint:\n        shot_authors = {authors[i] for i in shot_set if authors[i] is not None}\n        excluded_idx |= {i for i, a in enumerate(authors) if a in shot_authors}\n    target_pool = [w for i, y in enumerate(raw) if i not in excluded_idx\n                   for w in peak_windows(y, k=1)]\n    if not target_pool:\n        extra = (f\" ({len(excluded_idx) - len(shot_set)} extra recordings excluded for \"\n                 f\"sharing a recordist with a shot)\" if recordist_disjoint else \"\")\n        raise ValueError(f\"{sp}: no held-out recordings left after taking {n_shot} \"\n                          f\"shots{extra}\")\n\n    own_embs = perch_embed(support)\n    PROTO = l2n(own_embs).mean(0); PROTO /= np.linalg.norm(PROTO) + 1e-9\n\n    pool_sp = {k: v for k, v in SPECIES_CACHE.items() if k != sp}\n    ranked  = sorted(pool_sp, key=lambda o: -cos(pool_sp[o][\"emb\"], PROTO))\n    if decoy_mode == \"hard\":\n        DECOYS = ranked[:N_DECOYS]\n    elif decoy_mode == \"medium\":\n        DECOYS = ranked[N_DECOYS:2 * N_DECOYS]\n        if len(DECOYS) < N_DECOYS:\n            raise ValueError(f\"{sp}: only {len(ranked)} candidates ranked, need \"\n                              f\"{2*N_DECOYS} for decoy_mode='medium'\")\n    elif decoy_mode == \"random\":\n        # Seeded from this call's `seed`, offset so it never coincides with any other\n        # rng drawn from the same seed in this function (shot choice, calibration).\n        DECOYS = list(np.random.default_rng(90_000 + seed).choice(\n            ranked, size=N_DECOYS, replace=False))\n    else:\n        raise ValueError(f\"unknown decoy_mode: {decoy_mode!r} (want hard|medium|random)\")\n    decoy_sims  = [cos(pool_sp[o][\"emb\"], PROTO) for o in DECOYS]\n    decoy_pools = {o: _deep_decoy_pool(o) for o in DECOYS}\n    calib_neg_species = [o for o in ranked if o not in DECOYS]\n\n    # ---- domain-matched perch_clf (identical protocol to v1, per target) ----\n    neg_calls = [w for o in calib_neg_species for w in pool_sp[o][\"calls\"]]\n    rng_cal = np.random.default_rng(11 + seed)\n    X_parts, y_all = [], []\n    for bed in calib_beds:\n        y_bed = bed.copy(); spans, centres, labels = [], [], []\n        plan = [(1, support)] * 12 + [(0, neg_calls)] * 12\n        plan = [plan[i] for i in rng_cal.permutation(len(plan))]\n        for i, (lab, pl) in enumerate(plan):\n            c = place(y_bed, pl, rng_cal, spans, SNR_LEVELS[i % len(SNR_LEVELS)])\n            if c is None: continue\n            centres.append(c + float(rng_cal.uniform(-0.25, 0.25))); labels.append(lab)\n        for _ in range(8):\n            pos = int(rng_cal.integers(0, len(y_bed) - CLEN))\n            if any(pos < e and pos + CLEN > s for s, e in spans): continue\n            centres.append((pos + CLEN/2) / SR); labels.append(0)\n        X_parts.append(perch_embed_ctx(to32k(y_bed), centres)); y_all += labels\n\n    X = l2n(np.concatenate(X_parts)); yv = np.array(y_all)\n    clf = _make_verifier().fit(X, yv)\n\n    cov = _covariates(support, own_embs, decoy_sims)\n    cov.update(species=sp, n_recordings=len(files), n_shot=n_shot,\n               shot_strategy=strategy, n_heldout_calls=len(target_pool),\n               clf_train_acc=float(clf.score(X, yv)),\n               decoy_mode=decoy_mode,\n               recordist_disjoint=bool(recordist_disjoint),\n               n_distinct_recordists=(len(set(authors)) if recordist_disjoint else -1),\n               # provenance: a results CSV that does not say which Perch produced it\n               # is not comparable to any other run.\n               perch_version=PERCH_VERSION, perch_embed_dim=EMBED_DIM,\n               perch_peak_norm=(PERCH_PEAK_NORM if PERCH_PEAK_NORM is not None else 0.0),\n               perch_backend=PERCH_BACKEND,\n               verifier=VERIFIER)\n\n    if verbose:\n        rd_note = (f\"  distinct_recordists={cov['n_distinct_recordists']}\"\n                   if recordist_disjoint else \"\")\n        print(f\"  [{sp}] shots={strategy}({n_shot})  heldout={len(target_pool)} calls  \"\n              f\"decoys={DECOYS} ({decoy_mode})  top-sim={cov['decoy_sim_top1']:.3f}  \"\n              f\"stereotypy={cov['stereotypy']:.3f}  verifier={VERIFIER}{rd_note}\")\n\n    _TARGET_CACHE[ck] = dict(support=support, target_pool=target_pool, PROTO=PROTO,\n                             DECOYS=DECOYS, decoy_pools=decoy_pools, X=X, yv=yv,\n                             neg_calls=neg_calls,\n                             cov={k: v for k, v in cov.items() if k != \"clf_train_acc\"})\n    return dict(species=sp, support=support, target_pool=target_pool, PROTO=PROTO,\n                DECOYS=DECOYS, decoy_pools=decoy_pools, perch_clf=clf, covariates=cov,\n                neg_calls=neg_calls)\n\nprint(\"\\nCELL 1 v2 ready. build_target(sp) is the entry point; TARGET_SET is exported \"\n      \"for the encoder-bank exclusion in CELL 2.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"65eeec31","cell_type":"markdown","source":"## Cell 2 — Two-Stage Detection, Multi-Species\n\n`arguswala_v2.py`","metadata":{}},{"id":"577d15be","cell_type":"code","source":"# ===== CELL 2 v2 / TWO-STAGE across N species =====\n# Needs CELL 1 v2 (arguswala_setup_v2.py).\n#\n# WHAT CHANGED vs v1:\n#\n#  (1) THE BUG. v1 line 216 read:\n#          sps = [s for s in meta.primary_label.unique() if s != \"whbsho3\"]\n#      The target was hard-coded while CELL 1 defined it as a variable. The moment a\n#      second target is evaluated, that line puts the new target INTO the encoder's\n#      training bank -- the encoder is pre-trained on the species it must never have\n#      seen, every number inflates, and nothing crashes. Now excludes all of TARGET_SET,\n#      with an assertion so it can never regress silently.\n#\n#  (2) ONE ENCODER, MANY TARGETS. Legitimate because no target is in the bank, and it\n#      turns O(species x encoder-trainings) into O(1).\n#\n#  (3) PURE-NEGATIVE CONTROL (neg_pass). Nothing in v1 ever ran on a recording with zero\n#      target calls -- in deployment that is the common case, and false alarms per hour\n#      on clean audio is the first number a conservation biologist asks for.\n#\n#  (4) FEATURE TOGGLE. Liang et al. (DCASE 2024 baseline) Table III: PCEN gives the best\n#      PRECISION of any front-end tested (68.0%) while log-Mel gives the best F1 (63.7 vs\n#      60.0). v1 assumed PCEN. Now it is a measured choice.\n#\n#  (5) COVARIATES + per-species results written to CSV, so the difficulty-factor analysis\n#      Nolasco et al. attempted and failed can actually be run afterwards.\nimport numpy as np, os, glob, copy, time, librosa, torch, torch.nn as nn, torch.optim as optim\nimport pandas as pd\n\nSR, N_FFT, HOP, N_MELS = 22050, 1024, 256, 128   # identical to the DCASE 2024 baseline\n                                                 # (22.05kHz / 1024 / 256 / 128) -- keep it\n                                                 # that way, it is what makes a DCASE run cheap\nSEG_SEC = 1.0; SEG_FR = int(np.ceil(SEG_SEC*SR/HOP)); CLEN = int(SEG_SEC*SR); EMB = 128\nFT_STEPS, FT_LR, FT_BATCH = 250, 1e-4, 16\nN_INJECT, N_NEG, INJ_SNR = 45, 200, (-6, 12)\nCAND_FLOOR, STRIDE, IOU_THR = 0.10, max(1, SEG_FR//3), 0.30\nMAX_EVENT_SEC = 10.0\nCALLS_PER_BED = 12\n\nFEATURE = \"pcen\"          # 'pcen' | 'logmel'\nSEEDS   = (7, 8, 9)       # raise to >=10 once the timing measurement says it fits\nENCODER_SEEDS = (0, 1, 2) # >1 entry measures encoder-training variance (roadmap 3.4).\n                          # Widened 14 Aug 2026 after the first real run: whbsho3 stage-1\n                          # AUC came back 0.517 here vs. the earlier single-target baseline\n                          # of 0.459. This tests ONE candidate cause -- training-run\n                          # stochasticity on a FIXED bank (build_banks() always uses\n                          # seed=0 regardless of this tuple, so bank COMPOSITION doesn't\n                          # vary here). A separate single-target comparison run is still\n                          # the only way to test the other candidate cause (bank\n                          # composition differs because this run excludes 2 species from\n                          # the encoder bank, the original baseline only excluded 1).\n\n# 18 Aug 2026: a Kaggle GPU-quota cutoff killed a run mid-way through CELL 4 after\n# CELL 2's 5.5-hour species loop had already completed and its CSV was safely\n# downloaded. Restarting the whole notebook to reach CELL 4 again would re-pay that\n# 5.5 hours for a result already in hand. SKIP_MAIN_LOOP=True runs bank-building and\n# encoder training (~10-20 min, everything CELL 4 actually needs) and stops there --\n# `res`/`wide`/the printed report below never run, and never need to. CELL 4's own\n# checkpointing (RESUME=True) then picks up the ladder at the first unfinished config.\n# Leave False for a normal full run.\nSKIP_MAIN_LOOP = False\n\n# ---------------------------------------------------------------- ablation ladder\n# Stage-1 rungs. Both default OFF so rung 0 reproduces the original pipeline exactly.\n# (Stage-2 rung lives in CELL 1 as VERIFIER = 'logreg' | 'proto'.)\nUSE_TIMEFILTER = False    # rung 2: TimeFilterAug on the support during fine-tuning\nUSE_HARD_NEG   = False    # rung 3: hard-negative mining in the fine-tuning loop\nHARD_NEG_KEEP  = 0.25     # fraction of candidate negatives retained as \"hard\"\n\n# rung 5 (CELL 5, confound-closing run 2 -- roadmap \"owed\" list): stage-1's per-target\n# fine-tuning has always drawn its negatives as arbitrary random crops of the bed\n# (whatever silence/wind/distant noise happens to be there) -- never a labelled OTHER\n# SPECIES call. The stage-2 Perch verifier already trains against real other-species\n# calls (calib_neg_species in build_target); stage 1 never has. Default OFF reproduces\n# every previously-published number byte-for-byte. When True, stage1() injects real\n# calls from ctx['neg_calls'] (the SAME non-decoy species pool the verifier uses) into a\n# copy of the clean recording and crops THOSE positions as negatives instead of arbitrary\n# background -- so both stages are being asked to reject the same kind of confuser.\nUSE_SPECIES_NEG = False\n\n# CELL 5, confound-closing run 1: how far a sliding Perch-embedding + verifier scan\n# steps between windows when scoring a recording with NO stage-1 CNN in the loop at all\n# (detect_perch_only below). Perch inference is the pipeline's dominant per-call cost, so\n# this trades runtime against localisation precision; narrow it only after the first\n# timing print says the wider stride is affordable.\nPERCH_ALONE_STRIDE_S = 0.5\n\ndef timefilteraug(x, rng, m_range=(3, 6), lo_db=-6.0, hi_db=8.0):\n    \"\"\"TimeFilterAug -- Zou et al. 2024 (DCASE 2023 Task 5, 1st place, 63.8% F), eq. 6.\n\n    Motivation, in their words: the five shots \"are typically clear, whereas the query set\n    for predictions often contains interference noise, predominantly from far-field sound\n    and background impulse noise.\" This simulates that by applying a random piecewise-\n    LINEAR gain envelope across TIME (not a mask, and not across frequency):\n\n      partition the window into m in [3,6] segments; draw a breakpoint gain g_i ~ U[0,1]\n      per boundary; within each segment interpolate linearly between adjacent g's; map\n      that ramp onto [lo_db, hi_db] = [-6, +8] dB; apply.\n\n    Their front-end is PCEN at 22050 Hz / n_fft 1024 / hop 256 -- identical to ARGUS, so\n    the dB bounds transfer directly rather than needing rescaling.\n    \"\"\"\n    T = x.shape[1]\n    m = int(rng.integers(m_range[0], m_range[1] + 1))\n    if T < m + 1:\n        return x\n    cuts  = np.sort(rng.choice(np.arange(1, T), size=m - 1, replace=False))\n    edges = np.concatenate([[0], cuts, [T]])\n    g = rng.random(m + 1)\n    alpha = np.concatenate([\n        np.linspace(g[i], g[i + 1], int(edges[i + 1] - edges[i]), endpoint=False)\n        for i in range(m)\n    ])\n    beta = (lo_db + (hi_db - lo_db) * alpha).astype(\"float32\")   # per-frame gain in dB\n    if FEATURE == \"logmel\":\n        # amplitude_to_db output is ALREADY logarithmic -> a dB gain is additive.\n        return x + beta[np.newaxis, :]\n    # PCEN is a linear-domain magnitude -> a dB gain is multiplicative.\n    return x * (10.0 ** (beta / 20.0))[np.newaxis, :].astype(\"float32\")\n\ndef aug_pos(x, rng):\n    \"\"\"Augmentation applied to positives (the injected support) during fine-tuning.\"\"\"\n    x = specaug(x, rng)\n    if USE_TIMEFILTER:\n        x = timefilteraug(x, rng)\n    return x\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\"); print(\"device:\", device)\n\ndef rms(x): return float(np.sqrt(np.mean(np.square(x, dtype=np.float64)) + 1e-12))\ndef ld(p, dur=None):\n    y, _ = librosa.load(p, sr=SR, mono=True, duration=dur); return y.astype(\"float32\")\n\nFEATURES = (\"pcen\", \"logmel\")   # every front-end an encoder is trained for\n\ndef feat(y, feature=None):\n    \"\"\"Front-end. PCEN vs log-Mel is an experiment, not an assumption.\n\n    `feature=None` reads the FEATURE global -- that is what makes the ladder able to\n    switch front-ends by rebinding a global. Pass it explicitly when featurising for a\n    front-end other than the currently-selected one (e.g. building both encoder banks).\n    \"\"\"\n    feature = feature or FEATURE\n    m = librosa.feature.melspectrogram(y=y, sr=SR, n_fft=N_FFT, hop_length=HOP,\n                                       n_mels=N_MELS, power=1.0)\n    if feature == \"pcen\":\n        return librosa.pcen(m * (2**31), sr=SR, hop_length=HOP).astype(\"float32\")\n    if feature == \"logmel\":\n        return librosa.amplitude_to_db(m, ref=np.max).astype(\"float32\")\n    raise ValueError(f\"unknown feature: {feature}\")\n\ndef crop(f, c):\n    a = c - SEG_FR//2; p = np.full((N_MELS, SEG_FR), f.min(), dtype=\"float32\")\n    lo, hi = max(0, a), min(f.shape[1], a+SEG_FR)\n    if hi > lo: p[:, (lo-a):(lo-a)+(hi-lo)] = f[:, lo:hi]\n    return p\ndef loud_off(y, n=CLEN):\n    if len(y) <= n: return 0\n    e = np.square(y, dtype=np.float64); cs = np.concatenate([[0.0], np.cumsum(e)])\n    return int(np.argmax(cs[n:]-cs[:-n]))\ndef specaug(x, rng):\n    if rng.random() >= 0.6: return x\n    x = x.copy(); fl = x.min()\n    for _ in range(2):\n        k = int(rng.integers(0, 21))\n        if k: a = int(rng.integers(0, max(1, N_MELS-k))); x[a:a+k, :] = fl\n        k = int(rng.integers(0, 11))\n        if k: a = int(rng.integers(0, max(1, SEG_FR-k))); x[:, a:a+k] = fl\n    return x\ndef inject(bed, calls, n, snr, rng, gap_s=0.4):\n    y = bed.astype(\"float32\").copy(); L = len(y); gap = int(gap_s*SR); sp = []\n    if L <= CLEN: return y, sp\n    for _ in range(n*60):\n        if len(sp) >= n: break\n        pos = int(rng.integers(0, L-CLEN))\n        if any(pos < e+gap and pos+CLEN > s-gap for s, e in sp): continue\n        c = calls[int(rng.integers(len(calls)))][:CLEN]\n        if len(c) < CLEN: c = np.pad(c, (0, CLEN-len(c)))\n        g = rms(y[pos:pos+CLEN])*(10**(float(rng.uniform(*snr))/20))/(rms(c)+1e-9)\n        y[pos:pos+CLEN] += c.astype(\"float32\")*g; sp.append((pos, pos+CLEN))\n    return y, sorted(sp)\ndef iou(a, b):\n    it = max(0.0, min(a[1], b[1]) - max(a[0], b[0]))\n    u = (a[1]-a[0]) + (b[1]-b[0]) - it\n    return it/u if u > 0 else 0.0\n\nclass Enc(nn.Module):\n    def __init__(s, emb=EMB):\n        super().__init__()\n        def B(i, o): return nn.Sequential(nn.Conv2d(i, o, 3, padding=1, bias=False),\n                                          nn.BatchNorm2d(o), nn.ReLU(), nn.MaxPool2d(2))\n        s.e = nn.Sequential(B(1,128), B(128,128), B(128,128), B(128,emb)); s.p = nn.AdaptiveAvgPool2d(1)\n    def forward(s, x): return s.p(s.e(x.unsqueeze(1))).flatten(1)\n\ndef train_enc(bank, rng, episodes=1000, nway=6, k=4, q=4):\n    m = Enc().to(device); opt = optim.Adam(m.parameters(), 1e-3); ce = nn.CrossEntropyLoss()\n    cl = list(bank); m.train()\n    for ep in range(episodes):\n        ch = rng.choice(cl, size=min(nway, len(cl)), replace=False)\n        sx, sy, qx, qy = [], [], [], []\n        for lab, c in enumerate(ch):\n            idx = rng.permutation(len(bank[c]))[:k+q]\n            for j, i in enumerate(idx):\n                (sx if j < k else qx).append(torch.tensor(specaug(bank[c][i], rng)))\n                (sy if j < k else qy).append(lab)\n        sx = torch.stack(sx).to(device); qx = torch.stack(qx).to(device)\n        sy = torch.tensor(sy).to(device); qy = torch.tensor(qy).to(device)\n        se, qe = m(sx), m(qx)\n        pr = torch.stack([se[sy == c].mean(0) for c in torch.unique(sy)])\n        loss = ce(-torch.cdist(qe, pr), qy); opt.zero_grad(); loss.backward(); opt.step()\n        if (ep+1) % 250 == 0: print(f\"    enc ep {ep+1}/{episodes} loss {loss.item():.3f}\")\n    return m\n\ndef stitch(frames, probs):\n    ab = probs >= CAND_FLOOR; ev = []; i = 0\n    while i < len(frames):\n        if not ab[i]: i += 1; continue\n        j = i\n        while j+1 < len(frames) and ab[j+1] and frames[j+1]-frames[j] <= STRIDE*1.5: j += 1\n        pk = i + int(np.argmax(probs[i:j+1]))\n        s = max(0.0, frames[i]*HOP/SR - SEG_SEC/2); e = frames[j]*HOP/SR + SEG_SEC/2\n        if e - s > MAX_EVENT_SEC:\n            c = frames[pk]*HOP/SR; s, e = max(0.0, c-SEG_SEC/2), c+SEG_SEC/2\n        ev.append((s, e, float(probs[pk]))); i = j+1\n    return ev\n\ndef stitch_time(times, probs, floor=CAND_FLOOR, gap_s=None, max_event_sec=MAX_EVENT_SEC,\n                 seg_sec=SEG_SEC):\n    \"\"\"Same merge-adjacent-above-floor logic as stitch(), but for detectors that never\n    build a mel spectrogram (CELL 5's Perch-alone-as-detector), so positions are already\n    in seconds rather than STRIDE-spaced spectrogram frame indices. `times` must be sorted\n    ascending. gap_s defaults to a spacing consistent with stitch()'s STRIDE*1.5 in frames,\n    converted to seconds.\"\"\"\n    gap_s = gap_s if gap_s is not None else (STRIDE * HOP / SR) * 1.5\n    times = np.asarray(times, dtype=\"float64\")\n    ab = np.asarray(probs) >= floor; ev = []; i = 0\n    while i < len(times):\n        if not ab[i]: i += 1; continue\n        j = i\n        while j+1 < len(times) and ab[j+1] and times[j+1]-times[j] <= gap_s: j += 1\n        pk = i + int(np.argmax(probs[i:j+1]))\n        s = max(0.0, times[i] - seg_sec/2); e = times[j] + seg_sec/2\n        if e - s > max_event_sec:\n            c = times[pk]; s, e = max(0.0, c - seg_sec/2), c + seg_sec/2\n        ev.append((s, e, float(probs[pk]))); i = j+1\n    return ev\n\ndef stage1(enc, rec, sup, rng, neg_species_calls=None):\n    n_inj = max(8, int(round(N_INJECT * len(rec) / (240.0 * SR))))\n    aw, sp = inject(rec, sup, n_inj, INJ_SNR, rng)\n    af, sf = feat(aw), feat(rec)\n    pc = [((a+b)//2)//HOP for a, b in sp]\n    if not pc: return []\n    pos = [crop(af, c) for c in pc]\n    g = SEG_FR; hi = af.shape[1]-g\n\n    def _background_neg():\n        out = []\n        for _ in range(N_NEG*40):\n            if len(out) >= N_NEG: break\n            f = int(rng.integers(g, max(g+1, hi)))\n            if all(abs(f-p) > SEG_FR for p in pc): out.append(crop(af, f))\n        return out\n\n    if USE_SPECIES_NEG and neg_species_calls:\n        # Inject real OTHER-SPECIES calls into a FRESH copy of the clean recording (not\n        # `aw`, which already carries the positive injections -- a separate copy means\n        # positive and negative spans can never collide or corrupt each other) and crop\n        # those exact centres as negatives, instead of arbitrary background.\n        nw, nsp = inject(rec, neg_species_calls, N_NEG, INJ_SNR, rng)\n        nf = feat(nw)\n        nc = [((a+b)//2)//HOP for a, b in nsp]\n        neg = [crop(nf, c) for c in nc if g <= c <= nf.shape[1]-g]\n        if len(neg) < N_NEG:\n            # short recording / crowded bed couldn't place enough species-matched calls --\n            # top up with background rather than silently training on fewer negatives.\n            neg += _background_neg()\n    else:\n        neg = _background_neg()\n    if not neg:\n        neg = [np.full((N_MELS, SEG_FR), af.min(), dtype=\"float32\")]\n    m = copy.deepcopy(enc).to(device); h = nn.Linear(EMB, 2).to(device)\n    opt = optim.Adam(list(m.parameters())+list(h.parameters()), FT_LR)\n    ce = nn.CrossEntropyLoss(); m.train(); h.train(); hf = FT_BATCH//2\n\n    def _train(steps, neg_pool):\n        for _ in range(steps):\n            xs  = [aug_pos(pos[int(rng.integers(len(pos)))], rng) for _ in range(hf)]\n            xs += [neg_pool[int(rng.integers(len(neg_pool)))] for _ in range(hf)]\n            xb = torch.tensor(np.stack(xs)).to(device)\n            yb = torch.tensor([1]*hf+[0]*hf).to(device)\n            loss = ce(h(m(xb)), yb); opt.zero_grad(); loss.backward(); opt.step()\n\n    def _score(items):\n        m.eval(); h.eval(); out = []\n        with torch.no_grad():\n            for i in range(0, len(items), 256):\n                b = torch.tensor(np.stack(items[i:i+256])).to(device)\n                out.append(torch.softmax(h(m(b)), 1)[:, 1].cpu().numpy())\n        m.train(); h.train()\n        return np.concatenate(out)\n\n    if USE_HARD_NEG and len(neg) >= 8:\n        # Negative choice is named as decisive twice in the literature: Nolasco et al.\n        # sec.4.2 (\"prototype-based meta-learning works well when taking care about ...\n        # the choice of negative examples\") and Liang et al., who measure +5.31 F1 from\n        # negative hard sampling. Random background windows are mostly trivial negatives\n        # -- silence and wind -- so half the budget is spent learning nothing. Train\n        # briefly, find the background the model currently CONFUSES with the target, and\n        # spend the rest of the budget there.\n        half = FT_STEPS // 2\n        _train(half, neg)\n        k = max(4, int(round(len(neg) * HARD_NEG_KEEP)))\n        hard = [neg[i] for i in np.argsort(-_score(neg))[:k]]\n        _train(FT_STEPS - half, hard)\n    else:\n        _train(FT_STEPS, neg)\n    m.eval(); h.eval()\n    fr = list(range(g, max(g+1, sf.shape[1]-g), STRIDE)); pb = np.empty(len(fr), dtype=\"float32\")\n    with torch.no_grad():\n        for i in range(0, len(fr), 256):\n            b = torch.tensor(np.stack([crop(sf, f) for f in fr[i:i+256]])).to(device)\n            p = torch.softmax(h(m(b)), 1)[:, 1]; pb[i:i+len(p)] = p.cpu().numpy()\n    return stitch(fr, pb)\n\ndef detect_pv(enc, ctx, rec, rng):\n    \"\"\"-> [(start, end, p, v)]. Target context is now explicit, not a global.\"\"\"\n    ev = stage1(enc, rec, ctx[\"support\"], rng, neg_species_calls=ctx.get(\"neg_calls\"))\n    if not ev: return []\n    centres = [(s + e) / 2 for s, e, _ in ev]\n    emb = l2n(perch_embed_ctx(to32k(rec), centres))\n    v = ctx[\"perch_clf\"].predict_proba(emb)[:, 1]\n    return [(s, e, p, float(vi)) for (s, e, p), vi in zip(ev, v)]\n\ndef detect_perch_only(enc, ctx, rec, rng, stride_s=None):\n    \"\"\"-> [(start, end, p, v)], p := v. CELL 5, confound-closing run 1: \"what did the CNN\n    contribute?\" No stage-1 encoder, no mel front-end, no fine-tuning loop anywhere in\n    this path -- Perch embeddings are computed directly on a sliding window across the\n    raw recording and scored with the SAME calibrated verifier (ctx['perch_clf']) every\n    other detector in this file uses, then merged into events with stitch_time(). `enc`\n    and `rng` are accepted only for call-site parity with detect_pv (so eval_pass /\n    spec_pass / neg_pass can swap detectors via one `detect_fn` parameter) -- this path\n    is deterministic given `rec`, so both are unused. p is set equal to v so the FUSIONS\n    machinery still works mechanically, but only the \"perch only (v)\" fusion is a\n    meaningful score here -- any p-dependent fusion rule applied to this detector's\n    output is measuring nothing, since p carries no independent information.\"\"\"\n    stride_s = stride_s if stride_s is not None else PERCH_ALONE_STRIDE_S\n    L = len(rec) / SR\n    times = ([L / 2.0] if L <= SEG_SEC else\n             list(np.arange(SEG_SEC / 2.0, L - SEG_SEC / 2.0 + 1e-9, stride_s)))\n    if not times: return []\n    emb = l2n(perch_embed_ctx(to32k(rec), times))\n    v = ctx[\"perch_clf\"].predict_proba(emb)[:, 1]\n    ev = stitch_time(np.asarray(times), v)\n    return [(s, e, vi, vi) for s, e, vi in ev]\n\nFUSIONS = {\n    \"stage-1 only  (p)\":  lambda p, v: p,\n    \"perch only    (v)\":  lambda p, v: v,\n    \"p * v\":              lambda p, v: p * v,\n    \"p * sqrt(v)\":        lambda p, v: p * np.sqrt(v),\n    \"sqrt(p) * v\":        lambda p, v: np.sqrt(p) * v,\n    \"geo mean sqrt(p*v)\": lambda p, v: np.sqrt(p * v),\n}\n\n# ------------------------------------------------------------------- metrics\ndef match(events, truth, iou_thresh):\n    pairs = sorted(((iou(e[:2], t), i, j) for i, e in enumerate(events)\n                    for j, t in enumerate(truth)), reverse=True)\n    used_e, used_t = set(), set()\n    for score, i, j in pairs:\n        if score <= iou_thresh: break\n        if i in used_e or j in used_t: continue\n        used_e.add(i); used_t.add(j)\n    return [(e[2], i in used_e) for i, e in enumerate(events)]\n\ndef match_truth(events, truth, iou_thresh):\n    pairs = sorted(((iou(e[:2], t), i, j) for i, e in enumerate(events)\n                    for j, t in enumerate(truth)), reverse=True)\n    ue, ut, out = set(), set(), [0.0] * len(truth)\n    for score, i, j in pairs:\n        if score <= iou_thresh: break\n        if i in ue or j in ut: continue\n        ue.add(i); ut.add(j); out[j] = events[i][2]\n    return out\n\ndef fp_per_tp(scored, n_truth, target_recall=0.5):\n    arr = sorted(scored, key=lambda x: -x[0]); tp = fp = 0\n    for score, hit in arr:\n        if hit: tp += 1\n        else:   fp += 1\n        if tp / n_truth >= target_recall:\n            return fp / tp if tp else float('inf')\n    return float('inf')\n\ndef summarise(scored, n_truth):\n    if not scored or not n_truth:\n        return dict(ap=0.0, f1=0.0, p=0.0, r=0.0, thr=1.0, thr_r90=1.0, n=len(scored))\n    arr = sorted(scored, key=lambda x: -x[0])\n    tp = fp = 0; ap = 0.0; prev_r = 0.0; best = (0.0, 0.0, 0.0, 1.0); thr_r90 = None\n    for score, hit in arr:\n        if hit: tp += 1\n        else:   fp += 1\n        pr = tp / (tp + fp); rc = tp / n_truth\n        ap += pr * (rc - prev_r); prev_r = rc\n        f1 = 2*pr*rc/(pr+rc) if pr + rc else 0.0\n        if f1 > best[0]: best = (f1, pr, rc, score)\n        # Recall-first operating point (N4): the HIGHEST threshold still achieving >=90%\n        # recall, i.e. best precision subject to a recall floor. Scores descend, so the\n        # first crossing is that threshold. Setting it to the score just BEFORE the\n        # crossing (the earlier bug) yields a point that never actually reaches 90%.\n        if thr_r90 is None and rc >= 0.90: thr_r90 = score\n    if thr_r90 is None: thr_r90 = arr[-1][0]   # 90% unreachable -> loosest available\n    return dict(ap=ap, f1=best[0], p=best[1], r=best[2], thr=best[3],\n                thr_r90=thr_r90, n=len(scored))\n\ndef _auc(pos, neg):\n    \"\"\"P(random target scores above random decoy), ties 0.5. Threshold-FREE.\n    Same metric family BirdCLEF 2024 adopted as its official score (macro ROC-AUC):\n    0.50 = cannot tell target from decoy, 1.00 = perfect species discrimination.\"\"\"\n    pos, neg = np.asarray(pos, float), np.asarray(neg, float)\n    if not len(pos) or not len(neg): return float('nan')\n    gt = float((pos[:, None] > neg[None, :]).sum())\n    eq = float((pos[:, None] == neg[None, :]).sum())\n    return (gt + 0.5 * eq) / (len(pos) * len(neg))\n\n# ------------------------------------------------------------------- passes\n# detect_fn is swappable (default detect_pv) so CELL 5 can run the identical planting /\n# matching logic against a different detector -- e.g. detect_perch_only -- and get a\n# directly comparable number rather than a separately-implemented, possibly-diverging one.\ndef eval_pass(enc, ctx, seed, detect_fn=detect_pv):\n    r = np.random.default_rng(seed); cache = []\n    for bed in eval_beds:\n        y = bed.copy(); L = len(y); spans, planted, snrs = [], [], []\n        for k in range(CALLS_PER_BED):\n            snr = SNR_LEVELS[k % len(SNR_LEVELS)]\n            placed = None\n            for _ in range(80):\n                pos = int(r.integers(0, L-CLEN))\n                if any(pos < e+int(0.4*SR) and pos+CLEN > s-int(0.4*SR) for s, e in spans): continue\n                placed = pos; break\n            if placed is None: continue\n            c = ctx[\"target_pool\"][int(r.integers(len(ctx[\"target_pool\"])))][:CLEN]\n            if len(c) < CLEN: c = np.pad(c, (0, CLEN-len(c)))\n            g = rms(y[placed:placed+CLEN])*(10**(snr/20))/(rms(c)+1e-9)\n            y[placed:placed+CLEN] += c.astype(\"float32\")*g\n            spans.append((placed, placed+CLEN))\n            planted.append((placed/SR, (placed+CLEN)/SR)); snrs.append(snr)\n        cache.append((detect_fn(enc, ctx, y, r), planted, snrs))\n    return cache\n\ndef spec_pass(enc, ctx, seed, snr_db=6.0, detect_fn=detect_pv):\n    \"\"\"Specificity control. snr_db is now a parameter: v1 fixed it at +6 dB, so species\n    discrimination had only ever been measured on loud, easy calls (roadmap 3.2).\"\"\"\n    r = np.random.default_rng(seed); cache = []\n    for bed in eval_beds:\n        y = bed.copy(); L = len(y); spans, items = [], []\n        plan = [(\"target\", ctx[\"target_pool\"])]*6\n        for n_, p_ in ctx[\"decoy_pools\"].items(): plan += [(n_, p_)]*3\n        for kind, pool in plan:\n            placed = None\n            for _ in range(80):\n                pos = int(r.integers(0, L-CLEN))\n                if any(pos < e+int(0.4*SR) and pos+CLEN > s-int(0.4*SR) for s, e in spans): continue\n                placed = pos; break\n            if placed is None or not pool: continue\n            c = pool[int(r.integers(len(pool)))][:CLEN]\n            if len(c) < CLEN: c = np.pad(c, (0, CLEN-len(c)))\n            g = rms(y[placed:placed+CLEN])*(10**(snr_db/20))/(rms(c)+1e-9)\n            y[placed:placed+CLEN] += c.astype(\"float32\")*g\n            spans.append((placed, placed+CLEN))\n            items.append((kind, (placed/SR, (placed+CLEN)/SR)))\n        cache.append((detect_fn(enc, ctx, y, r), items))\n    return cache\n\ndef neg_pass(enc, ctx, seed, detect_fn=detect_pv):\n    \"\"\"PURE NEGATIVE control -- unmodified beds, zero target calls planted.\n    Every detection here is a false alarm by construction. This is the deployment-\n    relevant number (false alarms per hour of clean audio) and v1 never measured it.\"\"\"\n    r = np.random.default_rng(seed + 500); cache = []\n    for bed in eval_beds:\n        cache.append(detect_fn(enc, ctx, bed.copy(), r))\n    return cache\n\n# ------------------------------------------------------------------- scoring\ndef score_eval(fuse, cache):\n    scored, n_truth, by_snr = [], 0, []\n    for dets, planted, snrs in cache:\n        ev = [(s, e, float(fuse(p, v))) for s, e, p, v in dets]\n        scored += match(ev, planted, IOU_THR); n_truth += len(planted)\n        by_snr += list(zip(snrs, match_truth(ev, planted, IOU_THR)))\n    m = summarise(scored, n_truth)\n    m[\"fp_per_tp@R50\"] = fp_per_tp(scored, n_truth, 0.5)\n    m[\"recall_by_snr\"] = {s: float(np.mean([sc >= m[\"thr\"] for ss, sc in by_snr if ss == s]))\n                          for s in SNR_LEVELS}\n    return m\n\ndef score_spec(fuse, thr, cache, decoy_names):\n    tg, dc = [], {k: [] for k in decoy_names}\n    for dets, items in cache:\n        ev = [(s, e, float(fuse(p, v))) for s, e, p, v in dets]\n        for kind, span in items:\n            best = max((iou(e[:2], span), e[2]) for e in ev) if ev else (0.0, 0.0)\n            sc = best[1] if best[0] > IOU_THR else 0.0\n            (tg if kind == \"target\" else dc[kind]).append(sc)\n    t = float(np.mean([s >= thr for s in tg])) if tg else 0.0\n    d = {k: float(np.mean([s >= thr for s in v])) for k, v in dc.items() if v}\n    md = float(np.mean(list(d.values()))) if d else 0.0\n    all_dc = [s for v in dc.values() for s in v]\n    return t, md, (md/t if t else float('nan')), d, _auc(tg, all_dc)\n\ndef spec_coverage(cache):\n    \"\"\"-> fraction of planted items stage 1 LOCALISED at all (IoU > IOU_THR), by kind.\n\n    Only the SNR sweep (roadmap 3.2) needs this, and there it is not optional. score_spec\n    assigns 0.0 to any planted item that was never localised, and _auc counts a 0.0-vs-0.0\n    pair as a tie worth 0.5. So an AUC drifting toward 0.50 at low SNR is ambiguous between\n    two different findings: \"target and decoy became indistinguishable\" and \"nothing was\n    detected, so everything tied\". Coverage separates them. Reporting the SNR curve without\n    it would invite exactly the misreading the curve is meant to settle.\n\n    Fusion-independent by construction -- a fusion rule changes an event's score, never its\n    span -- so this is computed once per cache rather than once per fusion rule.\n    \"\"\"\n    n = {\"target\": 0, \"decoy\": 0}; hit = {\"target\": 0, \"decoy\": 0}\n    for dets, items in cache:\n        for kind, span in items:\n            k = \"target\" if kind == \"target\" else \"decoy\"\n            n[k] += 1\n            best = max((iou((s, e), span) for s, e, _p, _v in dets), default=0.0)\n            hit[k] += int(best > IOU_THR)\n    return {k: (hit[k]/n[k] if n[k] else float('nan')) for k in n}\n\ndef score_spec_detected(fuse, cache):\n    \"\"\"-> (AUC over LOCALISED items only, n_target_localised, n_decoy_localised).\n\n    The companion to spec_coverage, and the direct measurement it was a proxy for.\n    score_spec's AUC mixes two effects that pull in the same direction at low SNR: genuine\n    species confusion, and stage 1 failing to localise the call at all -- which forces a\n    0.0 and drags the AUC toward 0.50 for a reason that has nothing to do with species\n    identity. Excluding unlocalised items (rather than scoring them 0.0) separates the two,\n    so \"specificity degraded\" and \"detection degraded\" can be told apart from the numbers\n    instead of argued about afterwards.\n\n    Returns nan when either side is empty -- at low enough SNR that is the honest answer,\n    and it must not be silently averaged in as if it were a measurement.\n    \"\"\"\n    tg, dc = [], []\n    for dets, items in cache:\n        ev = [(s, e, float(fuse(p, v))) for s, e, p, v in dets]\n        for kind, span in items:\n            best = max((iou(e[:2], span), e[2]) for e in ev) if ev else (0.0, 0.0)\n            if best[0] <= IOU_THR: continue      # never localised -> excluded, NOT zeroed\n            (tg if kind == \"target\" else dc).append(best[1])\n    return _auc(tg, dc), len(tg), len(dc)\n\ndef score_neg(fuse, thr, cache):\n    \"\"\"-> false alarms per hour on audio containing no target calls.\"\"\"\n    fa = sum(sum(1 for s, e, p, v in dets if float(fuse(p, v)) >= thr) for dets in cache)\n    hours = (len(cache) * BED_SEC) / 3600.0\n    return fa / hours if hours else float('nan')\n\n# ------------------------------- stage-1 encoder (ALL targets excluded -- the bug fix)\ndef build_banks(exclude, features=FEATURES, seed=0, max_species=40):\n    \"\"\"One bank per front-end, sharing a single pass over the audio.\n\n    An encoder trained on PCEN cannot be evaluated on log-Mel inputs, so the front-end\n    cannot be a same-encoder ladder rung -- it needs its own encoder. But the expensive\n    part is decoding audio, not featurising it, so each recording is loaded ONCE and\n    featurised for every front-end. Cost is one extra encoder training, not a second\n    pass over the dataset.\n    \"\"\"\n    rng = np.random.default_rng(seed)\n    sps = [s for s in meta.primary_label.unique() if s not in exclude]\n    rng.shuffle(sps)\n    banks = {f: {} for f in features}\n    for sp in sps:\n        if len(banks[features[0]]) >= max_species: break\n        segs = {f: [] for f in features}\n        for fn in meta[meta.primary_label == sp][\"filename\"].tolist()[:12]:\n            try:\n                y = ld(os.path.join(AUD, fn), 15)\n                c = loud_off(y)//HOP + SEG_FR//2\n                for f in features:\n                    segs[f].append(crop(feat(y, f), c))\n            except Exception:\n                pass\n        if len(segs[features[0]]) >= 8:\n            for f in features:\n                banks[f][sp] = segs[f]\n    return banks\n\nprint(f\"\\nbuilding encoder banks for {list(FEATURES)}, excluding all {len(TARGET_SET)} \"\n      f\"targets: {sorted(TARGET_SET)}\")\nbanks = build_banks(TARGET_SET, FEATURES, seed=0)\nfor _f, _b in banks.items():\n    assert not (TARGET_SET & set(_b)), \\\n        f\"LEAKAGE: target species present in {_f} encoder bank: {TARGET_SET & set(_b)}\"\n    print(f\"  {_f:<7} bank: {len(_b)} species, none of them targets (assertion passed)\")\n\n# Keyed by (feature, seed): a rung that switches the front-end must switch the encoder\n# with it, or it is measuring an encoder/feature mismatch rather than the front-end.\nencoders = {}\nfor _f in FEATURES:\n    for es in ENCODER_SEEDS:\n        print(f\"training stage-1 encoder (feature={_f}, seed={es})...\")\n        torch.manual_seed(es)\n        encoders[(_f, es)] = train_enc(banks[_f], np.random.default_rng(es))\n\nif SKIP_MAIN_LOOP:\n    print(\"SKIP_MAIN_LOOP=True -- stopping after CELL 2's encoder training.\")\n    print(\"res/wide/the per-species report do not exist this run. Proceed to CELL 4;\")\n    print(\"its own checkpointing resumes the ladder at the first unfinished config.\")\nelse:\n    # ------------------------------------------------------------------- main loop\n    rows, t_start = [], time.time()\n    for sp in TARGET_LIST:\n        print(f\"\\n{'='*78}\\nTARGET: {sp}\\n{'='*78}\")\n        t_sp = time.time()\n        try:\n            ctx = build_target(sp)\n        except Exception as ex:\n            print(f\"  SKIP {sp}: {type(ex).__name__}: {ex}\"); continue\n\n        for es in ENCODER_SEEDS:\n            enc = encoders[(FEATURE, es)]\n            per_seed = {name: [] for name in FUSIONS}\n            for sd in SEEDS:\n                t0 = time.time()\n                ec = eval_pass(enc, ctx, sd)\n                sc = spec_pass(enc, ctx, sd)\n                nc = neg_pass(enc, ctx, sd)\n                for name, fuse in FUSIONS.items():\n                    m = score_eval(fuse, ec)\n                    t, md, ratio, per, auc = score_spec(fuse, m[\"thr\"], sc, ctx[\"DECOYS\"])\n                    per_seed[name].append(dict(\n                        ap=m[\"ap\"], f1=m[\"f1\"], p=m[\"p\"], r=m[\"r\"],\n                        fptp=m[\"fp_per_tp@R50\"], auc=auc, ratio=ratio,\n                        fa_per_hr=score_neg(fuse, m[\"thr\"], nc),\n                        fa_per_hr_r90=score_neg(fuse, m[\"thr_r90\"], nc)))\n                print(f\"  seed {sd} done in {time.time()-t0:.0f}s\")\n\n            for name in FUSIONS:\n                agg = {k: float(np.mean([d[k] for d in per_seed[name]])) for k in per_seed[name][0]}\n                sds = {k + \"_sd\": float(np.std([d[k] for d in per_seed[name]], ddof=1))\n                       if len(SEEDS) > 1 else 0.0 for k in per_seed[name][0]}\n                rows.append({**ctx[\"covariates\"], \"encoder_seed\": es, \"fusion\": name,\n                             \"feature\": FEATURE, \"n_seeds\": len(SEEDS),\n                             \"use_timefilter\": USE_TIMEFILTER, \"use_hard_neg\": USE_HARD_NEG,\n                             **agg, **sds})\n        print(f\"  [{sp}] total {time.time()-t_sp:.0f}s\")\n\n    res = pd.DataFrame(rows)\n    res.to_csv(\"/kaggle/working/argus_multispecies_results.csv\", index=False)\n    print(f\"\\nwrote {len(res)} rows -> argus_multispecies_results.csv \"\n          f\"| wall clock {(time.time()-t_start)/60:.1f} min\")\n\n    # ------------------------------------------------------------------- headline table\n    BASE, CAND = \"stage-1 only  (p)\", \"perch only    (v)\"\n    print(f\"\\n{'='*100}\\nSPECIFICITY AUC BY SPECIES  (threshold-free; 0.50 = target \"\n          f\"indistinguishable from decoys)\\n{'='*100}\")\n    print(f\"{'species':<12}{'n_rec':>6}{'decoy_sim':>10}{'stereo':>8}\"\n          f\"{'AUC stage-1':>13}{'AUC verified':>14}{'delta':>8}{'F1 s1':>7}{'F1 ver':>8}\")\n    print(\"-\"*100)\n    # Average over encoder seeds first, so a multi-encoder-seed run reports the mean rather\n    # than silently taking whichever row happened to land first.\n    NUM = [\"auc\", \"f1\", \"fa_per_hr\", \"fa_per_hr_r90\", \"n_recordings\",\n           \"decoy_sim_top1\", \"stereotypy\"]\n    piv = (res[res.fusion.isin([BASE, CAND])]\n           .groupby([\"species\", \"fusion\"], sort=False)[NUM].mean().reset_index())\n    wide = piv.pivot(index=\"species\", columns=\"fusion\")\n\n    for sp in [s for s in TARGET_LIST if s in wide.index]:\n        r = wide.loc[sp]\n        print(f\"{sp:<12}{int(r[('n_recordings', BASE)]):>6}{r[('decoy_sim_top1', BASE)]:>10.3f}\"\n              f\"{r[('stereotypy', BASE)]:>8.3f}{r[('auc', BASE)]:>13.3f}{r[('auc', CAND)]:>14.3f}\"\n              f\"{r[('auc', CAND)] - r[('auc', BASE)]:>+8.3f}\"\n              f\"{r[('f1', BASE)]:>7.3f}{r[('f1', CAND)]:>8.3f}\")\n\n    # ---------------------------------------------------------- encoder-seed spread\n    # The table above averages over encoder seeds -- which hides exactly the variance\n    # ENCODER_SEEDS was widened to measure. Same bank, same data, same eval seeds; the only\n    # thing differing is the encoder training run. Anything smaller than this spread is not\n    # a real effect, no matter what a within-run SE says about it.\n    if len(ENCODER_SEEDS) > 1:\n        print(f\"\\n{'='*100}\\nENCODER-SEED SPREAD -- pure training-run variance \"\n              f\"(same bank, same data, {len(ENCODER_SEEDS)} encoder seeds)\\n{'='*100}\")\n        es_piv = (res[res.fusion.isin([BASE, CAND])]\n                  .groupby([\"species\", \"fusion\", \"encoder_seed\"])[\"auc\"].mean().reset_index())\n        worst = 0.0\n        for sp in [s for s in TARGET_LIST if s in set(es_piv.species)]:\n            for fus, tag in ((BASE, \"stage-1\"), (CAND, \"verified\")):\n                g = es_piv[(es_piv.species == sp) & (es_piv.fusion == fus)].sort_values(\"encoder_seed\")\n                if g.empty: continue\n                v = g.auc.to_numpy(); spread = float(v.max() - v.min())\n                worst = max(worst, spread)\n                print(f\"  {sp:<10} {tag:<9} \" + \"  \".join(f\"{x:.3f}\" for x in v)\n                      + f\"   | spread {spread:+.3f}  sd {v.std(ddof=1):.3f}\")\n        print(f\"\\n  Largest spread from encoder training alone: {worst:.3f}\")\n        print(f\"  Treat any AUC difference below ~{worst:.3f} as noise, not signal.\")\n\n    if len(wide) > 1:\n        d = (wide[(\"auc\", CAND)] - wide[(\"auc\", BASE)]).to_numpy()\n        se = d.std(ddof=1)/np.sqrt(len(d))\n        print(f\"\\nACROSS SPECIES (n={len(d)}): mean AUC gain {d.mean():+.3f}, SE {se:.3f}, \"\n              f\"delta/SE {d.mean()/se if se else float('nan'):.2f}   (|d/SE| < 2 is NOT evidence)\")\n        print(f\"  verification helped:            {int((d > 0).sum())}/{len(d)} species\")\n        print(f\"  below chance WITHOUT verifying: \"\n              f\"{int((wide[('auc', BASE)] < 0.5).sum())}/{len(d)} species\")\n        print(f\"  below chance WITH verifying:    \"\n              f\"{int((wide[('auc', CAND)] < 0.5).sum())}/{len(d)} species\")\n        # The pre-registered outcome test (roadmap sec.2): does the original single-species\n        # finding generalise, hold only conditionally, or fail? Reported either way.\n        frac = (d > 0).mean()\n        print(\"  -> outcome \" + (\"A (generalises)\" if frac >= 0.8 else\n                                 \"B (conditional -- find what separates them)\" if frac >= 0.4 else\n                                 \"C/D (does NOT generalise -- rebuild narrative around the boundary)\"))\n\n    print(f\"\\nFALSE ALARMS PER HOUR on clean audio (no target present)\")\n    print(f\"{'species':<12}{'best-F1 thr':>24}{'recall-first thr':>26}\")\n    for sp in [s for s in TARGET_LIST if s in wide.index]:\n        r = wide.loc[sp]\n        print(f\"  {sp:<10}{r[('fa_per_hr', BASE)]:>10.1f} ->{r[('fa_per_hr', CAND)]:>8.1f}\"\n              f\"{r[('fa_per_hr_r90', BASE)]:>16.1f} ->{r[('fa_per_hr_r90', CAND)]:>8.1f}\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"21f6624f","cell_type":"markdown","source":"## Cell 3 — Shot-Selection & Shot-Count Sweeps\n\n`arguswala_sweeps.py`","metadata":{}},{"id":"c0052698","cell_type":"code","source":"# ===== CELL 3 / SWEEPS -- shot selection, shot count, front-end =====\n# Run AFTER arguswala_v2.py (reuses its trained `encoders`, scoring fns and eval passes).\n#\n# Five experiments the literature says matter and nobody has published answers for on\n# this data:\n#\n#  N1  SHOT-SELECTION STRATEGY.  Nolasco et al. (2023) sec.4.1: unrepresentative support\n#      sets were the single failure cause their expert annotators raised most often --\n#      \"the five shots ... are not representative of the range of possible [calls]\".\n#      DCASE always takes the FIRST five. ARGUS v1 always took the highest-RATED five.\n#      Nobody has measured what that choice is worth. Four strategies, same everything else.\n#\n#  S2  SHOT COUNT.  ARGUS is a few-shot project that has never varied the number of shots.\n#      First thing a judge asks. Where the curve plateaus is a deployable field number.\n#\n#  S3  SPECIFICITY vs SNR.  Every specificity number this project has reported was\n#      measured at a fixed +6 dB -- loud, close calls. Roadmap 3.2 calls this a real hole,\n#      with a pre-registered prediction attached. A deployment cannot choose its SNR.\n#\n#  S4  DECOY DIFFICULTY.  Decoys have always been the top-N_DECOYS hardest confusers, on\n#      purpose -- never measured against the alternative. Roadmap 3.3.\n#\n#  N2  FRONT-END.  Liang et al. (DCASE 2024 baseline) Table III: PCEN has the best\n#      PRECISION of any front-end tested (68.0%), log-Mel the best F1 (63.7 vs 60.0).\n#      ARGUS uses PCEN. Confirm or contradict that trade on Western Ghats data.\n#      NOTE: changing FEATURE changes the encoder bank too, so this needs a fresh\n#      encoder -- it is NOT a free re-score. Budget one extra encoder training.\nimport numpy as np, pandas as pd, time, torch\n\nSWEEP_SPECIES = TARGET_LIST[:3]      # sweeps are O(strategies x species x seeds); keep narrow\nSWEEP_SEEDS   = (7, 8)               # fewer than the main run -- these are relative comparisons\nBASE, CAND    = \"stage-1 only  (p)\", \"perch only    (v)\"\nENC           = encoders[(FEATURE, ENCODER_SEEDS[0])]   # encoders are keyed by front-end\n\ndef run_one(sp, n_shot, strategy, seeds=SWEEP_SEEDS, decoy_mode=\"hard\"):\n    \"\"\"One (species, n_shot, strategy, decoy_mode) cell -> mean metrics over seeds.\n    decoy_mode defaults to 'hard' (roadmap 3.3's meaning of \"unchanged\") so N1/S2 above\n    are byte-for-byte the same run they always were; only S4 below passes anything else.\"\"\"\n    try:\n        ctx = build_target(sp, n_shot=n_shot, strategy=strategy, verbose=False,\n                            decoy_mode=decoy_mode)\n    except Exception as ex:\n        print(f\"    skip {sp} n_shot={n_shot} {strategy} {decoy_mode}: \"\n              f\"{type(ex).__name__}: {ex}\")\n        return None\n    acc = {BASE: [], CAND: []}\n    for sd in seeds:\n        ec = eval_pass(ENC, ctx, sd)\n        sc = spec_pass(ENC, ctx, sd)\n        for name in (BASE, CAND):\n            m = score_eval(FUSIONS[name], ec)\n            _, _, _, _, auc = score_spec(FUSIONS[name], m[\"thr\"], sc, ctx[\"DECOYS\"])\n            acc[name].append((auc, m[\"f1\"], m[\"p\"], m[\"r\"]))\n    out = dict(species=sp, n_shot=n_shot, strategy=strategy, decoy_mode=decoy_mode,\n               n_heldout=ctx[\"covariates\"][\"n_heldout_calls\"],\n               decoy_sim_top1=ctx[\"covariates\"][\"decoy_sim_top1\"])\n    for name, tag in ((BASE, \"s1\"), (CAND, \"ver\")):\n        a = np.array(acc[name], dtype=float)\n        out[f\"auc_{tag}\"] = a[:, 0].mean(); out[f\"auc_{tag}_sd\"] = a[:, 0].std(ddof=1) if len(a) > 1 else 0.0\n        out[f\"f1_{tag}\"]  = a[:, 1].mean()\n        out[f\"prec_{tag}\"] = a[:, 2].mean(); out[f\"rec_{tag}\"] = a[:, 3].mean()\n    return out\n\n# ------------------------------------------------------------------ N1: shot selection\nprint(f\"\\n{'='*84}\\nN1  SHOT-SELECTION STRATEGY  (n_shot={N_SHOT} held constant)\\n{'='*84}\")\nt0 = time.time(); rows_n1 = []\nfor strategy in (\"first\", \"rated\", \"random\", \"diverse\"):\n    for sp in SWEEP_SPECIES:\n        r = run_one(sp, N_SHOT, strategy)\n        if r: rows_n1.append(r); print(f\"  {strategy:<8} {sp:<12} \"\n                                       f\"AUC {r['auc_s1']:.3f} -> {r['auc_ver']:.3f}   \"\n                                       f\"F1 {r['f1_s1']:.3f} -> {r['f1_ver']:.3f}\")\nn1 = pd.DataFrame(rows_n1)\nif len(n1):\n    n1.to_csv(\"/kaggle/working/argus_sweep_shot_strategy.csv\", index=False)\n    print(f\"\\n{'strategy':<10}{'AUC stage-1':>13}{'AUC verified':>14}{'F1 stage-1':>12}{'F1 verified':>13}\")\n    print(\"-\"*62)\n    for s, g in n1.groupby(\"strategy\", sort=False):\n        print(f\"{s:<10}{g.auc_s1.mean():>13.3f}{g.auc_ver.mean():>14.3f}\"\n              f\"{g.f1_s1.mean():>12.3f}{g.f1_ver.mean():>13.3f}\")\n    best = n1.groupby(\"strategy\").auc_ver.mean().idxmax()\n    spread = n1.groupby(\"strategy\").auc_ver.mean()\n    print(f\"\\nbest strategy by verified AUC: {best}  \"\n          f\"(spread across strategies: {spread.max()-spread.min():.3f} AUC)\")\n    print(\"If that spread exceeds the verification gain itself, then HOW the five shots are\"\n          \"\\nchosen matters more than the architecture -- which is a publishable observation\"\n          \"\\nand exactly the open problem Nolasco et al. sec.4.1 raise.\")\nprint(f\"[N1 took {(time.time()-t0)/60:.1f} min]\")\n\n# ------------------------------------------------------------------ S2: shot count\nprint(f\"\\n{'='*84}\\nS2  SHOT COUNT  (strategy='{SHOT_STRATEGY}' held constant)\\n{'='*84}\")\nt0 = time.time(); rows_s2 = []\nfor n_shot in (1, 3, 5, 10):\n    for sp in SWEEP_SPECIES:\n        r = run_one(sp, n_shot, SHOT_STRATEGY)\n        if r: rows_s2.append(r); print(f\"  n_shot={n_shot:<3} {sp:<12} \"\n                                       f\"AUC {r['auc_s1']:.3f} -> {r['auc_ver']:.3f}   \"\n                                       f\"F1 {r['f1_s1']:.3f} -> {r['f1_ver']:.3f}\")\ns2 = pd.DataFrame(rows_s2)\nif len(s2):\n    s2.to_csv(\"/kaggle/working/argus_sweep_shot_count.csv\", index=False)\n    print(f\"\\n{'n_shot':<8}{'AUC stage-1':>13}{'AUC verified':>14}{'F1 verified':>13}\")\n    print(\"-\"*48)\n    for n, g in s2.groupby(\"n_shot\"):\n        print(f\"{n:<8}{g.auc_s1.mean():>13.3f}{g.auc_ver.mean():>14.3f}{g.f1_ver.mean():>13.3f}\")\n    print(\"\\nRead the plateau, not the peak: the smallest n_shot within noise of the best is\"\n          \"\\nthe number a field team actually has to collect.\")\nprint(f\"[S2 took {(time.time()-t0)/60:.1f} min]\")\n\n# ------------------------------------------------------------- S3: specificity vs SNR\n# Roadmap 3.2 -- the last unmeasured axis in the scenario matrix. spec_pass() has always\n# planted at a fixed +6 dB, so EVERY species-discrimination number this project has ever\n# reported -- including the 0.751 centrepiece -- describes loud, close calls only. A field\n# deployment does not get to choose the SNR.\n#\n# PRE-REGISTERED PREDICTION (roadmap 3.2, written before this ever ran):\n#   \"the verifier's advantage shrinks as calls get fainter.\"\n#   If true  -> an honest limitation with a number attached, and the centrepiece result\n#               must be quoted as a HIGH-SNR result from then on.\n#   If false -> a stronger claim than the project currently makes.\n# Both outcomes get reported. Nothing about the verdict is decided after seeing the curve.\nSNR_SWEEP_SPECIES = SWEEP_SPECIES     # same 3 as N1/S2, so all three sweeps stay comparable\nSNR_SWEEP_SEEDS   = SWEEP_SEEDS\nSNR_SWEEP_LEVELS  = SNR_LEVELS        # (9, 6, 3, 0, -3, -6); trim first if budget is tight\nNOISE_FLOOR       = 0.025             # measured in-session Day 26-27. Nothing below it is real.\n\ndef run_snr(sp, seeds=SNR_SWEEP_SEEDS, levels=SNR_SWEEP_LEVELS):\n    \"\"\"One species -> one row per (seed, SNR level, fusion).\n\n    eval_pass runs ONCE per seed rather than once per level. It exists here only to set the\n    operating threshold, and its own planting protocol is SNR-independent by design (it\n    rotates through SNR_LEVELS internally). Re-running it per level would cost 6x for an\n    identical threshold -- and would also be the wrong experiment. Calibrating once and then\n    varying the call level IS the deployment case: you fix a detector's operating point, and\n    the forest hands you whatever it hands you.\n    \"\"\"\n    try:\n        ctx = build_target(sp, verbose=False)\n    except Exception as ex:\n        print(f\"    skip {sp}: {type(ex).__name__}: {ex}\")\n        return []\n    rows = []\n    for sd in seeds:\n        ec  = eval_pass(ENC, ctx, sd)\n        thr = {name: score_eval(FUSIONS[name], ec)[\"thr\"] for name in (BASE, CAND)}\n        for snr in levels:\n            sc  = spec_pass(ENC, ctx, sd, snr_db=float(snr))\n            cov = spec_coverage(sc)          # fusion-independent, so computed once per level\n            for name, tag in ((BASE, \"s1\"), (CAND, \"ver\")):\n                t, md, ratio, _per, auc = score_spec(FUSIONS[name], thr[name], sc, ctx[\"DECOYS\"])\n                auc_det, n_td, n_dd = score_spec_detected(FUSIONS[name], sc)\n                rows.append(dict(species=sp, seed=sd, snr_db=int(snr), fusion=tag,\n                                 auc=auc, auc_det=auc_det, n_tgt_det=n_td, n_dec_det=n_dd,\n                                 tgt_rate=t, decoy_rate=md, ratio=ratio,\n                                 cov_target=cov[\"target\"], cov_decoy=cov[\"decoy\"],\n                                 thr=float(thr[name])))\n        print(f\"    {sp:<12} seed {sd} done ({len(levels)} levels)\")\n    return rows\n\n_np = len(SNR_SWEEP_SPECIES) * len(SNR_SWEEP_SEEDS) * (1 + len(SNR_SWEEP_LEVELS))\nprint(f\"\\n{'='*84}\\nS3  SPECIFICITY vs SNR  (roadmap 3.2)\\n{'='*84}\")\nprint(f\"{len(SNR_SWEEP_SPECIES)} species x {len(SNR_SWEEP_SEEDS)} seeds x \"\n      f\"({len(SNR_SWEEP_LEVELS)} levels + 1 threshold pass) = {_np} detection passes \"\n      f\"~= {_np*1.1:.0f} min at the measured 1.1 min/pass.\")\nprint(\"PRE-REGISTERED: the verification gain is predicted to SHRINK as SNR falls.\")\nprint(\"Reported either way -- a flat curve is the stronger result, not a failed run.\\n\")\n\nt0 = time.time(); rows_s3 = []\nfor sp in SNR_SWEEP_SPECIES:\n    rows_s3 += run_snr(sp)\ns3 = pd.DataFrame(rows_s3)\n\nif len(s3):\n    s3.to_csv(\"/kaggle/working/argus_sweep_snr.csv\", index=False)\n    print(f\"\\n{'SNR':>6}{'AUC s1':>9}{'AUC ver':>9}{'gain':>8}{'sd':>7}\"\n          f\"{'AUCdet s1':>11}{'AUCdet ver':>12}{'cov tgt':>9}{'cov dec':>9}\")\n    print(\"-\"*80)\n    curve = {}\n    for snr in sorted(s3.snr_db.unique(), reverse=True):\n        d  = s3[s3.snr_db == snr]\n        s1 = d[d.fusion == \"s1\"]; ve = d[d.fusion == \"ver\"]\n        # coverage is fusion-independent, so either subset reports it\n        curve[int(snr)] = c = dict(\n            auc_s1=float(s1.auc.mean()), auc_ver=float(ve.auc.mean()),\n            gain=float(ve.auc.mean() - s1.auc.mean()),\n            det_s1=float(s1.auc_det.mean()), det_ver=float(ve.auc_det.mean()),\n            cov_t=float(ve.cov_target.mean()), cov_d=float(ve.cov_decoy.mean()),\n            sd=float(ve.auc.std(ddof=1)) if len(ve) > 1 else 0.0,\n            sd_det=float(ve.auc_det.std(ddof=1)) if len(ve) > 1 else 0.0)\n        print(f\"{snr:>+5d}dB{c['auc_s1']:>9.3f}{c['auc_ver']:>9.3f}{c['gain']:>+8.3f}\"\n              f\"{c['sd']:>7.3f}{c['det_s1']:>11.3f}{c['det_ver']:>12.3f}\"\n              f\"{c['cov_t']:>9.0%}{c['cov_d']:>9.0%}\")\n\n    COV_MIN = 0.75    # below this, raw AUC is contaminated by detection failure (see caveat)\n    KNIFE   = 0.25    # a margin under 25% of the bound is a coin flip, not a finding\n\n    def trend(label, keys, get, sd_key, note, flat_note):\n        \"\"\"One trend test, reporting the MARGIN explicitly.\n\n        A bare `drop > bound` turns a 0.001 difference into a printed 'CONFIRMED' -- which\n        is how this project has produced retractions before. The margin is printed every\n        time, and the verdict is THREE-WAY, not two:\n\n          margin >  +KNIFE*bound : |change| clears the noise -> a real trend\n          margin <  -KNIFE*bound : |change| sits far BELOW the noise -> confidently FLAT\n          otherwise              : |change| ~= the noise -> genuinely unresolved\n\n        The middle and bottom cases are opposites and must not be merged. An earlier\n        version of this function tested only `margin < KNIFE*bound` and so reported a\n        change of exactly 0.000 against a 0.045 bound -- the strongest possible evidence\n        of flatness -- as \"too close to call\", burying the actual result (23 Aug 2026).\n        \"\"\"\n        if len(keys) < 2:\n            print(f\"\\n{label}\\n  -> NO VERDICT: {len(keys)} level(s) is a point, not a trend.\")\n            return\n        v_hi, v_lo = get(curve[keys[0]]), get(curve[keys[-1]])\n        if not (v_hi == v_hi and v_lo == v_lo):          # nan guard\n            print(f\"\\n{label}\\n  -> NO VERDICT: not measurable at an endpoint (nan).\")\n            return\n        drop   = v_hi - v_lo\n        bound  = max(NOISE_FLOOR, max(curve[k][sd_key] for k in keys))\n        margin = abs(drop) - bound\n        print(f\"\\n{label}\")\n        print(f\"  {v_hi:+.3f} at {keys[0]:+d} dB -> {v_lo:+.3f} at {keys[-1]:+d} dB    \"\n              f\"change {drop:+.3f}   bound {bound:.3f}   margin {margin:+.4f}\")\n        if margin > KNIFE * bound:\n            if drop > 0:\n                print(f\"  -> SHRINKS by {drop:.3f}, clear of the noise bound. {note}\")\n            else:\n                print(f\"  -> GROWS by {-drop:.3f} as calls get fainter. Report as observed; do\"\n                      f\"\\n     not rationalise it into the write-up without a separate test.\")\n        elif margin < -KNIFE * bound:\n            print(f\"  -> FLAT within noise. |change| ({abs(drop):.3f}) sits well below the \"\n                  f\"noise bound\\n     ({bound:.3f}) -- this is positive evidence of no trend, \"\n                  f\"not an absent result. {flat_note}\")\n        else:\n            print(f\"  -> TOO CLOSE TO CALL. |change| and the noise bound differ by \"\n                  f\"{margin:+.4f}, within\\n     {KNIFE:.0%} of the bound itself -- a coin flip. \"\n                  f\"Unresolved: widen SNR_SWEEP_SEEDS\\n     before claiming flat OR a direction.\")\n\n    allk   = sorted(curve, reverse=True)\n    validk = sorted([k for k, c in curve.items() if c[\"cov_t\"] >= COV_MIN], reverse=True)\n\n    trend(\"[1] verification GAIN, raw AUC, full SNR range\", allk,\n          lambda c: c[\"gain\"], \"sd\",\n          \"Quote the centrepiece as a high-SNR result from here on.\",\n          \"The verification gain holds across the whole tested SNR range.\")\n    if validk != allk:\n        trend(f\"[2] verification GAIN, raw AUC, coverage>={COV_MIN:.0%} levels only\", validk,\n              lambda c: c[\"gain\"], \"sd\",\n              \"Holds where detection is still intact -- a real specificity effect.\",\n              \"Where detection is intact, the gain does not depend on SNR.\")\n    trend(\"[3] verified AUC among LOCALISED items only (detection-corrected)\", allk,\n          lambda c: c[\"det_ver\"], \"sd_det\",\n          \"Specificity itself degrades, independently of detection.\",\n          \"Species discrimination is SNR-INVARIANT once a call is localised -- so any fall\"\n          \"\\n     in raw AUC is detection loss, not species confusion. This is the strong result.\")\n    trend(\"[4] stage-1 AUC among LOCALISED items only (control)\", allk,\n          lambda c: c[\"det_s1\"], \"sd_det\",\n          \"Stage 1's own specificity degrades too.\",\n          \"Stage 1's own specificity is stable too, so [3] is not an artefact of the verifier.\")\n\n    print(f\"\\n{'-'*80}\\nHOW TO READ [1] AND [3] TOGETHER -- they answer different questions\")\n    print(\"  raw falls + detection-corrected FLAT  -> DETECTION degrades with SNR; species\")\n    print(\"                                           discrimination, given a detection, does not.\")\n    print(\"  raw falls + detection-corrected FALLS -> specificity genuinely degrades; the\")\n    print(\"                                           limitation is real and belongs in the paper.\")\n    print(\"  both flat                             -> robust across the whole tested range.\")\n    print(\"  [4] is the control: if stage-1's corrected curve moves while the verified one does\")\n    print(\"  not, the verifier is what stabilises it -- and that is the claim worth making.\")\n\n    lo = min(curve)\n    if curve[lo][\"cov_t\"] < COV_MIN:\n        bad = [k for k in allk if curve[k][\"cov_t\"] < COV_MIN]\n        print(f\"\\n  CAVEAT -- raw AUC at {bad} dB is CONTAMINATED. Target coverage there is \"\n              f\"{curve[lo]['cov_t']:.0%};\\n  score_spec scores an unlocalised item 0.0 and _auc \"\n              f\"counts 0.0-vs-0.0 as a tie, so those levels\\n  measure detection failure as much \"\n              f\"as species confusion. Worse, coverage is ASYMMETRIC\\n  (target \"\n              f\"{curve[lo]['cov_t']:.0%} vs decoy {curve[lo]['cov_d']:.0%}): unmatched zeros bias \"\n              f\"the AUC in a fixed\\n  direction rather than merely adding noise. Use the AUCdet \"\n              f\"columns at those levels, or\\n  restrict the claim to the coverage>={COV_MIN:.0%} \"\n              f\"range and say so explicitly.\")\nprint(f\"[S3 took {(time.time()-t0)/60:.1f} min]\")\n\n# ---------------------------------------------------------- S4: decoy difficulty sweep\n# Roadmap 3.3. Decoys have always been the top-N_DECOYS nearest species in Perch space --\n# the hardest possible confusers, chosen on purpose. That choice was never actually\n# measured against the alternative: how much of the headline number is because the decoys\n# are hard? This turns one number into a difficulty curve, and pre-empts \"you picked easy\n# decoys\" -- you didn't, but until this ran that was an assertion, not a result.\n#\n# PRE-REGISTERED, and split into what's actually at risk of being wrong:\n#   Not a real prediction (true by construction of how DECOYS is ranked): AUC should rise\n#   hard -> medium -> random, for both stage-1 and verified. Decoys get progressively less\n#   similar to the target by definition, so this ordering isn't a finding, it's a sanity\n#   check on the ranking itself.\n#   The actual open question: does VERIFICATION'S GAIN (ver - s1) shrink, grow, or stay\n#   flat as decoys get easier? No strong prior either way -- reported honestly, whichever\n#   way it lands.\nDECOY_SWEEP_SPECIES = SWEEP_SPECIES\nDECOY_SWEEP_SEEDS   = SWEEP_SEEDS\nDECOY_MODES         = (\"hard\", \"medium\", \"random\")\n\nprint(f\"\\n{'='*84}\\nS4  DECOY DIFFICULTY  (roadmap 3.3)\\n{'='*84}\")\nt0 = time.time(); rows_s4 = []\nfor mode in DECOY_MODES:\n    for sp in DECOY_SWEEP_SPECIES:\n        r = run_one(sp, N_SHOT, SHOT_STRATEGY, seeds=DECOY_SWEEP_SEEDS, decoy_mode=mode)\n        if r: rows_s4.append(r); print(f\"  {mode:<7} {sp:<12} decoy-sim {r['decoy_sim_top1']:.3f}  \"\n                                       f\"AUC {r['auc_s1']:.3f} -> {r['auc_ver']:.3f}   \"\n                                       f\"F1 {r['f1_s1']:.3f} -> {r['f1_ver']:.3f}\")\ns4 = pd.DataFrame(rows_s4)\n\nif len(s4):\n    s4.to_csv(\"/kaggle/working/argus_sweep_decoy_difficulty.csv\", index=False)\n    print(f\"\\n{'mode':<9}{'decoy-sim':>11}{'AUC stage-1':>13}{'AUC verified':>14}{'gain':>8}{'gain sd':>9}\")\n    print(\"-\"*64)\n    agg = {}\n    for mode, g in s4.groupby(\"decoy_mode\", sort=False):\n        gains = g.auc_ver - g.auc_s1\n        agg[mode] = dict(sim=g.decoy_sim_top1.mean(), s1=g.auc_s1.mean(), ver=g.auc_ver.mean(),\n                         gain=gains.mean(), gain_sd=gains.std(ddof=1) if len(gains) > 1 else 0.0)\n        c = agg[mode]\n        print(f\"{mode:<9}{c['sim']:>11.3f}{c['s1']:>13.3f}{c['ver']:>14.3f}{c['gain']:>+8.3f}\"\n              f\"{c['gain_sd']:>9.3f}\")\n\n    ordered = [agg[m] for m in DECOY_MODES if m in agg]\n    print(f\"\\nsanity checks -- judged on AUC, which is what 'harder' actually MEANS:\")\n    # AUC ordering is the real difficulty check. If the tiers are separating at all, a\n    # harder decoy set must score LOWER, for stage-1 and verified alike.\n    auc_ok_s1  = all(ordered[i][\"s1\"]  <= ordered[i+1][\"s1\"]  + 1e-9 for i in range(len(ordered)-1))\n    auc_ok_ver = all(ordered[i][\"ver\"] <= ordered[i+1][\"ver\"] + 1e-9 for i in range(len(ordered)-1))\n    print(f\"  AUC rises hard -> medium -> random, stage-1:  \"\n          f\"{'OK' if auc_ok_s1 else 'VIOLATED -- tiers are not separating, investigate'}\")\n    print(f\"  AUC rises hard -> medium -> random, verified: \"\n          f\"{'OK' if auc_ok_ver else 'VIOLATED -- tiers are not separating, investigate'}\")\n    # The one similarity relation that IS true by construction: 'hard' takes ranked[0],\n    # the literal maximum, so nothing can beat it on top-1.\n    if \"hard\" in agg:\n        hard_is_max = all(agg[\"hard\"][\"sim\"] >= agg[m][\"sim\"] - 1e-9 for m in agg)\n        print(f\"  'hard' holds the highest top-1 similarity (true by construction):  \"\n              f\"{'OK' if hard_is_max else 'VIOLATED -- decoy RANKING itself is broken'}\")\n    sim_str = \"   \".join(f\"{m}: sim {agg[m]['sim']:.3f}\" for m in DECOY_MODES if m in agg)\n    print(f\"\\n  {sim_str}\")\n    print(\"  NOTE: top-1 similarity does NOT fall hard > medium > random, and that is\")\n    print(\"  EXPECTED rather than a fault. 'medium' deliberately excludes ranks 1..N_DECOYS,\")\n    print(\"  so its top-1 is a deterministic ceiling at ranked[N_DECOYS]. 'random' samples the\")\n    print(\"  WHOLE pool -- with 30 candidates and 4 draws it catches one of the four hardest\")\n    print(\"  ~45% of the time -- and top-1 is a MAX over that draw, so it lands ABOVE medium's\")\n    print(\"  ceiling. 'random' is therefore average-difficulty-with-high-variance, not an easy\")\n    print(\"  tier. Read difficulty off the AUC columns, never off this scalar: the same lesson\")\n    print(\"  the decoy-pool confound already taught (Day 28) -- one similarity number does not\")\n    print(\"  capture a decoy SET's difficulty.\")\n\n    if \"hard\" in agg and \"random\" in agg:\n        drop = agg[\"random\"][\"gain\"] - agg[\"hard\"][\"gain\"]\n        bound = max(NOISE_FLOOR, agg[\"hard\"][\"gain_sd\"], agg[\"random\"][\"gain_sd\"])\n        margin = abs(drop) - bound\n        KNIFE = 0.25\n        print(f\"\\nTHE ACTUAL QUESTION -- does verification's gain depend on decoy difficulty?\")\n        print(f\"  gain on hardest decoys   {agg['hard']['gain']:+.3f}\")\n        print(f\"  gain on random decoys    {agg['random']['gain']:+.3f}\")\n        print(f\"  change {drop:+.3f}   bound {bound:.3f}   margin {margin:+.4f}\")\n        if margin > KNIFE * bound:\n            if drop > 0:\n                print(f\"  -> GAIN IS LARGER against easy decoys, clear of noise. Verification's\")\n                print(f\"     advantage is partly a property of how hard the decoys are, not just\")\n                print(f\"     the target species -- the 0.751-class headline number is a HARD-DECOY\")\n                print(f\"     number and should keep being described as one.\")\n            else:\n                print(f\"  -> GAIN IS LARGER against the hardest decoys, clear of noise. Verification\")\n                print(f\"     earns its keep precisely where it's needed most -- the headline number\")\n                print(f\"     is not being flattered by an easy comparison; if anything it's the\")\n                print(f\"     hardest case, understating what an easier deployment would see.\")\n        elif margin < -KNIFE * bound:\n            print(f\"  -> FLAT within noise. Verification's advantage does not depend on how hard\")\n            print(f\"     the decoys are -- the 0.751-class headline number generalises across the\")\n            print(f\"     difficulty curve, not just at the one operating point it was measured at.\")\n        else:\n            print(f\"  -> TOO CLOSE TO CALL, within {KNIFE:.0%} of the bound. Unresolved at this\")\n            print(f\"     seed count -- widen DECOY_SWEEP_SEEDS before claiming a direction.\")\n\n    print(f\"\\nWhy this matters beyond the sanity check: the headline 0.751-class AUC has always\")\n    print(f\"been reported on the HARDEST decoy set on purpose -- the number a skeptical judge\")\n    print(f\"would ask for. The random-decoy row above is the number if we'd taken the easy path\")\n    print(f\"instead, so the difference between them is now a stated fact, not an assertion.\")\nprint(f\"[S4 took {(time.time()-t0)/60:.1f} min]\")\n\n# ------------------------------------------------------------------ N7: Perch version\n# Two-run protocol: a kernel holds ONE Perch model, so this cannot be a re-score.\n# DO THIS BEFORE THE FULL MULTI-SPECIES RUN -- decoys are ranked by Perch embedding\n# proximity, so changing Perch changes the decoy sets and makes runs incomparable.\nprint(f\"\\n{'='*84}\\nN7  PERCH 1.0 vs 2.0 -- manual two-run protocol, DO THIS FIRST\\n{'='*84}\")\nprint(f\"\"\"Currently loaded: Perch {PERCH_VERSION}, {EMBED_DIM}-d embeddings, peak-norm {PERCH_PEAK_NORM}\n\nTo compare:\n  1. Run the whole notebook with PERCH_VERSION = \"v2\" (default), SMOKE_TEST = True\n     -> save argus_multispecies_results.csv as ..._perchv2.csv\n  2. Set PERCH_VERSION = \"v1\" in CELL 1 and attach the v1 Kaggle model variation\n     (bird-vocalization-classifier/tensorFlow2/bird-vocalization-classifier, version 8)\n  3. Restart the kernel and re-run -> save as ..._perchv1.csv\n  4. Compare verified AUC per species.\n\nPublished expectation (Perch 2.0 paper, Table 3): BirdSet AUROC 0.839 -> 0.908 for the\nsame embeddings, no fine-tuning. If ARGUS sees no gain, that is worth reporting -- it\nwould suggest the bottleneck is stage 1, not the embedding.\n\nNOTE the confound: peak normalisation was ALSO wrong in v1 (raw waveforms were fed to\nPerch; the official wrapper DC-removes and scales to peak 0.25). To attribute a change to\nthe model rather than the preprocessing, run v1 twice -- PERCH_PEAK_NORM = None and 0.25 --\nor accept that the comparison bundles both fixes and say so explicitly.\"\"\")\n\n# ------------------------------------------------------------------ N2: front-end\n# NO LONGER A MANUAL PROTOCOL. CELL 2 now trains one encoder per front-end in a single\n# pass over the audio (build_banks), so PCEN vs log-Mel is a proper factor in the\n# ablation ladder (CELL 4) rather than a two-run comparison across kernels.\nprint(f\"\\n{'='*84}\\nN2  FRONT-END (PCEN vs log-Mel) -- now handled by CELL 4\\n{'='*84}\")\nprint(\"\"\"Run arguswala_ladder.py. It appears as:\n  - cumulative \"rung 4  +log-Mel\", or\n  - a main effect in the 2^4 design with FULL_FACTORIAL = True (16 configs).\n\nThe ladder resolves encoders[(FEATURE, seed)] per config, so the encoder always matches\nthe front-end -- evaluating a PCEN-trained encoder on log-Mel inputs would measure a\nmismatch, not the front-end.\n\nExpected, and genuinely contested:\n  - DCASE 2024 official table, same team and architecture: log-Mel 65.2% vs PCEN 61.5%\n  - Liang et al. Table III: log-Mel 63.67 vs PCEN 59.97 F1 -- but PCEN wins precision\n    (68.0 vs 66.4), and species specificity is a precision-shaped problem\n  - DCASE 2023 winner (Zou et al.) used PCEN at exactly ARGUS's 22050/1024/256 settings\nTwo of three point to log-Mel on F1; the 2023 winner and the precision column point back\nat PCEN. This is why it is measured rather than assumed.\"\"\")\n\nprint(f\"\\n{'='*84}\\nALL SWEEPS DONE. Copy the three CSVs out of /kaggle/working/ and log the\"\n      f\"\\nheadline numbers in the bound logbook the same day.\\n{'='*84}\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"be74354c","cell_type":"markdown","source":"## Cell 4 — Ablation Ladder\n\n`arguswala_ladder.py`","metadata":{}},{"id":"0b75f362","cell_type":"code","source":"# ===== CELL 4 / ABLATION LADDER =====\n# Run AFTER arguswala_v2.py (reuses its trained `encoders`, passes and scorers).\n#\n# REVISED 14 Aug 2026 after the first factorial run was killed at config 6/16.\n# Three faults in the previous version, all fixed here:\n#\n#  (1) NO CHECKPOINTING. Results were only written to CSV after the whole loop finished,\n#      so a killed session lost every completed config. ~82 min of real compute was\n#      recoverable only from scrollback. Now each config is appended to CSV the moment\n#      it finishes, and a re-run resumes instead of restarting.\n#\n#  (2) CATASTROPHIC CONFIG ORDERING. itertools.product with VERIFIER first made it the\n#      slowest-varying factor -- all 8 logreg configs ran before any proto config. The\n#      run died at 6/16, so it produced ZERO data on the prototypical probe, which was\n#      the entire reason for running a factorial. Configs are now ordered by DISTANCE\n#      FROM BASELINE, so the first 5 cover the baseline plus every single-factor main\n#      effect. A truncated run now still answers the main question.\n#\n#  (3) NO NOISE FLOOR. The previous run accidentally measured one: `pcen|logreg` scored\n#      0.696 in the cumulative run and 0.731 in the factorial run -- identical config,\n#      identical seeds, 0.035 apart, purely from GPU non-determinism (cuDNN conv backward\n#      is not bit-reproducible even with a fixed seed). Without that number, a delta of\n#      -0.005 was reported as \"HURTS (d/SE -3.02)\" when it is in fact indistinguishable\n#      from zero. The baseline is now deliberately repeated across encoder seeds to\n#      measure the floor, and every delta is judged against it.\nimport numpy as np, pandas as pd, time, itertools, os, gc\n\nLADDER_SPECIES = TARGET_LIST[:3]     # widen once the timing measurement says it fits\nLADDER_SEEDS   = (7, 8)\nFULL_FACTORIAL = True\n\n# TimeFilterAug is EXCLUDED by default as of 14 Aug 2026. Measured in isolation\n# (hn=False, both front-ends) it hurt every metric: verified AUC -0.155 (pcen) and\n# -0.088 (logmel), F1 0.713->0.522 and 0.483->0.362, and stage-1-alone AUC 0.493->0.407\n# -- that last one is verifier-independent, so the damage is unambiguously in stage-1\n# fine-tuning, not an interaction with the probe. At 4.4x the measured noise floor this\n# is a real effect, not scatter. Dropping it halves the design (16 -> 8 configs).\n# Set False to put it back and re-test.\nSKIP_TIMEFILTER = True\n\n# The baseline is re-run once per encoder seed to establish the run-to-run noise floor.\n# Deltas smaller than this are not evidence, regardless of what a within-run SE says.\nBASELINE_ENCODER_SEEDS = ENCODER_SEEDS          # from CELL 2; (0,1,2) after the widening\nOTHER_ENCODER_SEED     = ENCODER_SEEDS[0]\n\nCKPT_PATH = \"/kaggle/working/argus_ablation_ladder.csv\"\nRESUME    = True     # skip (config, encoder_seed) pairs already present in CKPT_PATH\n\n# 24 Aug 2026: the ladder finished a complete 8-config run (all 10 planned rows) that day\n# -- see the lab notebook Day 29 entry. RESUME's checkpoint only helps within ONE Kaggle\n# session (/kaggle/working/ is wiped between sessions unless manually re-attached), so a\n# fresh session re-derives an answer this project already has rather than resuming past\n# it. SKIP_LADDER=True skips this cell's run loop entirely when you only need CELL 3\n# (e.g. just the decoy-difficulty sweep, S4) and the ladder's own conclusion isn't what\n# you're re-running for. Leave False for a normal run, or when re-verifying the ladder is\n# actually the point (e.g. re-testing at a wider species count).\nSKIP_LADDER = False\n\nBASE, CAND = \"stage-1 only  (p)\", \"perch only    (v)\"\nBASELINE_CFG = dict(VERIFIER=\"logreg\", USE_TIMEFILTER=False, USE_HARD_NEG=False, FEATURE=\"pcen\")\n\nCUMULATIVE = [\n    (\"rung 0  baseline\",        dict(VERIFIER=\"logreg\", USE_TIMEFILTER=False, USE_HARD_NEG=False, FEATURE=\"pcen\")),\n    (\"rung 1  +proto probe\",    dict(VERIFIER=\"proto\",  USE_TIMEFILTER=False, USE_HARD_NEG=False, FEATURE=\"pcen\")),\n    (\"rung 2  +timefilter\",     dict(VERIFIER=\"proto\",  USE_TIMEFILTER=True,  USE_HARD_NEG=False, FEATURE=\"pcen\")),\n    (\"rung 3  +hard negatives\", dict(VERIFIER=\"proto\",  USE_TIMEFILTER=True,  USE_HARD_NEG=True,  FEATURE=\"pcen\")),\n    (\"rung 4  +log-Mel\",        dict(VERIFIER=\"proto\",  USE_TIMEFILTER=True,  USE_HARD_NEG=True,  FEATURE=\"logmel\")),\n]\n\ndef _tag(cfg):\n    t = f\"{cfg['FEATURE']}|{cfg['VERIFIER']}\"\n    if cfg[\"USE_TIMEFILTER\"]: t += \"+tfa\"\n    if cfg[\"USE_HARD_NEG\"]:   t += \"+hn\"\n    return t\n\ndef _dist_from_baseline(cfg):\n    return sum(1 for k, v in BASELINE_CFG.items() if cfg[k] != v)\n\ndef _factorial():\n    tfa_levels = (False,) if SKIP_TIMEFILTER else (False, True)\n    out = []\n    for v, t, hn, ft in itertools.product((\"logreg\", \"proto\"), tfa_levels,\n                                          (False, True), (\"pcen\", \"logmel\")):\n        cfg = dict(VERIFIER=v, USE_TIMEFILTER=t, USE_HARD_NEG=hn, FEATURE=ft)\n        out.append((_tag(cfg), cfg))\n    # Order by how many factors differ from baseline. With 4 binary factors this yields\n    # 1 baseline, then all single-factor changes, then pairs, then triples. A run killed\n    # partway still delivers every main effect -- which is what the previous ordering\n    # failed to do.\n    out.sort(key=lambda kc: (_dist_from_baseline(kc[1]), kc[0]))\n    return out\n\nCONFIGS = _factorial() if FULL_FACTORIAL else CUMULATIVE\nBASELINE_LABEL = _tag(BASELINE_CFG)\n\ndef _apply(cfg):\n    \"\"\"CELL 1 and CELL 2 share one notebook namespace, and _make_verifier(), aug_pos(),\n    stage1() and feat() all read these as globals at CALL time -- so rebinding here is\n    what switches the config. Verified by asserting the read-back below.\"\"\"\n    g = globals()\n    for k, val in cfg.items():\n        g[k] = val\n    assert (VERIFIER, USE_TIMEFILTER, USE_HARD_NEG, FEATURE) == \\\n           (cfg[\"VERIFIER\"], cfg[\"USE_TIMEFILTER\"], cfg[\"USE_HARD_NEG\"], cfg[\"FEATURE\"]), \\\n           \"flag rebind failed\"\n\ndef run_config(label, cfg, enc_seed):\n    _apply(cfg)\n    # The encoder MUST match the front-end: one trained on PCEN cannot be evaluated on\n    # log-Mel inputs, and doing so measures a mismatch rather than the front-end.\n    key = (FEATURE, enc_seed)\n    if key not in encoders:\n        raise KeyError(f\"no encoder for {key}; add '{FEATURE}' to FEATURES in CELL 2 \"\n                       f\"and re-run the encoder training. Have: {sorted(encoders)}\")\n    ENC = encoders[key]\n    rows = []\n    for sp in LADDER_SPECIES:\n        try:\n            ctx = build_target(sp, verbose=False)     # cached; only the verifier refits\n        except Exception as ex:\n            print(f\"    skip {sp}: {type(ex).__name__}: {ex}\"); continue\n        acc = {BASE: [], CAND: []}\n        for sd in LADDER_SEEDS:\n            ec = eval_pass(ENC, ctx, sd)\n            sc = spec_pass(ENC, ctx, sd)\n            nc = neg_pass(ENC, ctx, sd)\n            for name in (BASE, CAND):\n                mm = score_eval(FUSIONS[name], ec)\n                _, _, _, _, auc = score_spec(FUSIONS[name], mm[\"thr\"], sc, ctx[\"DECOYS\"])\n                acc[name].append((auc, mm[\"f1\"], mm[\"p\"], mm[\"r\"],\n                                  score_neg(FUSIONS[name], mm[\"thr\"], nc)))\n        row = dict(config=label, species=sp, encoder_seed=enc_seed, **{k: cfg[k] for k in cfg})\n        for name, tag in ((BASE, \"s1\"), (CAND, \"ver\")):\n            a = np.array(acc[name], dtype=float)\n            row[f\"auc_{tag}\"]  = a[:, 0].mean()\n            row[f\"f1_{tag}\"]   = a[:, 1].mean()\n            row[f\"prec_{tag}\"] = a[:, 2].mean()\n            row[f\"rec_{tag}\"]  = a[:, 3].mean()\n            row[f\"fa_{tag}\"]   = a[:, 4].mean()\n        rows.append(row)\n    return rows\n\n# ------------------------------------------------------------------ build the run plan\n# Baseline gets every encoder seed (noise floor); everything else gets one.\nPLAN = []\nfor label, cfg in CONFIGS:\n    seeds = BASELINE_ENCODER_SEEDS if label == BASELINE_LABEL else (OTHER_ENCODER_SEED,)\n    for es in seeds:\n        PLAN.append((label, cfg, es))\n\ndone = set()\nall_rows = []\nif RESUME and os.path.exists(CKPT_PATH):\n    prev = pd.read_csv(CKPT_PATH)\n    if {\"config\", \"encoder_seed\"}.issubset(prev.columns):\n        all_rows = prev.to_dict(\"records\")\n        done = {(r[\"config\"], int(r[\"encoder_seed\"])) for r in all_rows}\n        print(f\"RESUMING: {len(done)} (config, encoder_seed) pairs already in {CKPT_PATH}\")\n\ntodo = [(l, c, e) for (l, c, e) in PLAN if (l, int(e)) not in done]\nprint(f\"\\n{'='*100}\\nABLATION LADDER  |  {len(PLAN)} runs planned \"\n      f\"({len(CONFIGS)} configs, baseline x{len(BASELINE_ENCODER_SEEDS)} for noise floor) \"\n      f\"| {len(todo)} to run\\n\"\n      f\"  species={list(LADDER_SPECIES)}  eval seeds={LADDER_SEEDS}  \"\n      f\"timefilter={'EXCLUDED' if SKIP_TIMEFILTER else 'included'}\\n\"\n      f\"  checkpointing to {CKPT_PATH} after EVERY config -- a killed session resumes, \"\n      f\"it does not restart.\\n{'='*100}\")\n\nif SKIP_LADDER:\n    print(\"SKIP_LADDER=True -- skipping the ladder's run loop entirely.\")\n    print(\"This cell already has a complete answer from a prior session (see the\")\n    print(\"lab notebook Day 29 entry); re-running it here would just re-derive it.\")\nelse:\n    t0 = time.time()\n    for i, (label, cfg, es) in enumerate(todo, 1):\n        t1 = time.time()\n        r = run_config(label, cfg, es)\n        all_rows += r\n        # Write immediately. This is the whole point -- 82 min of compute was lost last run.\n        pd.DataFrame(all_rows).to_csv(CKPT_PATH, index=False)\n        if r:\n            d = pd.DataFrame(r)\n            print(f\"[{i:2d}/{len(todo)}] {label:<22} enc{es}  \"\n                  f\"AUC {d.auc_s1.mean():.3f} -> {d.auc_ver.mean():.3f}   \"\n                  f\"F1 {d.f1_ver.mean():.3f}   FA/h {d.fa_ver.mean():5.1f}   \"\n                  f\"[{time.time()-t1:.0f}s]  saved\")\n        gc.collect()\n        try:\n            import torch; torch.cuda.empty_cache()\n        except Exception:\n            pass\n\n    lad = pd.DataFrame(all_rows)\n    print(f\"\\ntotal {(time.time()-t0)/60:.1f} min | {len(lad)} rows -> {CKPT_PATH}\")\n\n    # ------------------------------------------------------------------ noise floor first\n    noise = None\n    b = lad[lad.config == BASELINE_LABEL]\n    if not b.empty and b.encoder_seed.nunique() > 1:\n        per_seed = b.groupby(\"encoder_seed\").auc_ver.mean()\n        noise = float(per_seed.max() - per_seed.min())\n        print(f\"\\n{'='*100}\\nNOISE FLOOR -- baseline '{BASELINE_LABEL}' repeated across \"\n              f\"{per_seed.size} encoder seeds\\n{'='*100}\")\n        print(\"  verified AUC per encoder seed: \" + \"  \".join(f\"{v:.3f}\" for v in per_seed))\n        print(f\"  spread {noise:+.3f}   sd {per_seed.std(ddof=1):.3f}\")\n        print(\"  -> any delta below this spread is NOT evidence, whatever a within-run SE says.\")\n        print(\"  (Cross-session noise is larger still: the same config scored 0.696 and 0.731 \"\n              \"in two\\n   separate sessions -- 0.035 apart -- from GPU non-determinism alone.)\")\n\n    # ------------------------------------------------------------------ report\n    print(f\"\\n{'='*100}\\n{'config':<24}{'enc':>4}{'AUC stage-1':>13}{'AUC verified':>14}\"\n          f\"{'delta vs base':>15}{'F1 ver':>9}{'FA/hr':>8}\\n{'-'*100}\")\n    base_auc = b.auc_ver.mean() if not b.empty else np.nan\n    for label, cfg in CONFIGS:\n        g = lad[lad.config == label]\n        if g.empty: continue\n        a = g.auc_ver.mean()\n        d = a - base_auc\n        flag = \"\"\n        if noise and label != BASELINE_LABEL:\n            flag = \"  <- within noise\" if abs(d) < noise else (\"  <- REAL\" if d < 0 else \"  <- REAL +\")\n        print(f\"{label:<24}{g.encoder_seed.nunique():>4}{g.auc_s1.mean():>13.3f}{a:>14.3f}\"\n              f\"{d:>+15.3f}{g.f1_ver.mean():>9.3f}{g.fa_ver.mean():>8.1f}{flag}\")\n\n    # ------------------------------------------------------------------ main effects\n    if FULL_FACTORIAL and not lad.empty:\n        print(f\"\\nMAIN EFFECTS (technique ON minus OFF, paired within species, \"\n              f\"averaged over all other settings):\")\n        for flag, on_vals, off_vals in ((\"VERIFIER\", [\"proto\"], [\"logreg\"]),\n                                        (\"USE_TIMEFILTER\", [True], [False]),\n                                        (\"USE_HARD_NEG\", [True], [False]),\n                                        (\"FEATURE\", [\"logmel\"], [\"pcen\"])):\n            if flag not in lad.columns: continue\n            g_on, g_off = lad[lad[flag].isin(on_vals)], lad[lad[flag].isin(off_vals)]\n            if g_on.empty or g_off.empty:\n                print(f\"  {flag:<34} not tested in this run\"); continue\n            on_s, off_s = g_on.groupby(\"species\").auc_ver.mean(), g_off.groupby(\"species\").auc_ver.mean()\n            common = on_s.index.intersection(off_s.index)\n            if len(common) < 2: continue\n            d = (on_s.loc[common] - off_s.loc[common]).to_numpy()\n            se = d.std(ddof=1) / np.sqrt(len(d)) if len(d) > 1 else float(\"nan\")\n            verdict = \"\"\n            if noise:\n                verdict = \"  within noise floor\" if abs(d.mean()) < noise else \"  exceeds noise floor\"\n            print(f\"  {flag} ({on_vals[0]} vs {off_vals[0]}):\".ljust(36)\n                  + f\"{d.mean():+.3f}  SE {se:.3f}  d/SE {d.mean()/se if se else float('nan'):>6.2f}{verdict}\")\n        print(\"  NOTE: within-run SE ignores GPU non-determinism between runs. Trust the \"\n              \"noise floor over d/SE.\")\n\n    print(\"\\nLog the noise floor and the surviving effects in the bound logbook the same day.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"0fd9d19e","cell_type":"markdown","source":"## Cell 5 — Confound-Closing Runs\n\n`arguswala_confounds.py`","metadata":{}},{"id":"bbec488f","cell_type":"code","source":"# ===== CELL 5 / CONFOUND-CLOSING RUNS =====\n# Run AFTER arguswala_v2.py (reuses its trained `encoders`, passes and scorers). Does NOT\n# depend on CELL 3 or CELL 4 -- like them, it only needs CELL 1+2, so it can run in any\n# order relative to the sweeps/ladder.\n#\n# Three items the 25 Aug 2026 adversarial review named as genuinely unrun (lab notebook,\n# Day 31): \"None of these are required before Oct 3; they close INTERVIEW risk, not\n# submission risk\" -- a judge is expected to ask each of these questions live.\n#\n#  P1  PERCH ALONE AS A STANDALONE DETECTOR. Every number this project has ever reported\n#      uses stage-1 (the trained CNN) to propose candidate windows and Perch only to\n#      VERIFY them. Nobody has asked: if you skip stage-1 entirely and just slide Perch +\n#      the calibrated verifier across the recording, what do you get? Answers \"what did\n#      the CNN actually contribute\" before a judge asks it live.\n#\n#  P2  STAGE-1 FINE-TUNED WITH SPECIES-MATCHED NEGATIVES. Stage-2's Perch verifier has\n#      always trained against real other-species calls as negatives (calib_neg_species in\n#      build_target). Stage-1's per-target fine-tuning (stage1() in CELL 2) never has --\n#      its negatives have always been arbitrary random crops of whatever's in the bed.\n#      USE_SPECIES_NEG (CELL 2) tests whether that choice matters.\n#\n#  P3  RECORDIST-DISJOINT HELD-OUT SPLIT. \"Did it learn the microphone, not the bird?\"\n#      The held-out pool has never been checked against WHO recorded the 5 support shots.\n#      recordist_disjoint=True (build_target, CELL 1) additionally excludes any held-out\n#      recording sharing a recordist with a shot.\nimport numpy as np, pandas as pd, time, os, gc\n\nCONFOUND_SPECIES = TARGET_LIST[:3]     # same narrowing convention as CELL 3/4 -- these\n                                       # are relative comparisons, not the headline number\nCONFOUND_SEEDS   = (7, 8)\nNOISE_FLOOR      = 0.025               # same measured floor CELL 3/4 use\nKNIFE            = 0.25\nBASE, CAND       = \"stage-1 only  (p)\", \"perch only    (v)\"\nENC              = encoders[(FEATURE, ENCODER_SEEDS[0])]\n\nCKPT_PATH = \"/kaggle/working/argus_confounds.csv\"\nRESUME    = True     # skip (run_type, species) pairs already checkpointed\n\n# ------------------------------------------------------------------ shared plumbing\ndef _mean_sd(vals):\n    a = np.asarray(vals, dtype=float)\n    return float(a.mean()), (float(a.std(ddof=1)) if len(a) > 1 else 0.0)\n\ndef _arm_metrics(fuse, ec, sc, nc, ctx):\n    m = score_eval(fuse, ec)\n    _, _, _, _, auc = score_spec(fuse, m[\"thr\"], sc, ctx[\"DECOYS\"])\n    fa = score_neg(fuse, m[\"thr\"], nc)\n    return dict(auc=auc, f1=m[\"f1\"], prec=m[\"p\"], rec=m[\"r\"], fa_per_hr=fa)\n\ndef verdict(label, val_a, val_b, sd_a, sd_b, pos_note, neg_note, flat_note,\n            noise_floor=NOISE_FLOOR, knife=KNIFE):\n    \"\"\"Two-arm version of CELL 3's trend() margin test -- same three-way logic (real /\n    flat / unresolved), just for an A-vs-B comparison instead of a multi-point curve. A\n    bare `b - a > 0` would turn a 0.001 difference into a printed verdict, which is\n    exactly the mistake CELL 3's own trend() was written to stop making.\"\"\"\n    drop = val_b - val_a\n    bound = max(noise_floor, sd_a, sd_b)\n    margin = abs(drop) - bound\n    print(f\"\\n{label}\")\n    print(f\"  {val_a:+.3f} -> {val_b:+.3f}   change {drop:+.3f}   bound {bound:.3f}   \"\n          f\"margin {margin:+.4f}\")\n    if margin > knife * bound:\n        print(f\"  -> {'RISES' if drop > 0 else 'FALLS'} by {abs(drop):.3f}, clear of the \"\n              f\"noise bound. {pos_note if drop > 0 else neg_note}\")\n    elif margin < -knife * bound:\n        print(f\"  -> FLAT within noise ({abs(drop):.3f} vs bound {bound:.3f}). {flat_note}\")\n    else:\n        print(f\"  -> TOO CLOSE TO CALL, within {knife:.0%} of the bound. Unresolved at \"\n              f\"this seed count -- widen CONFOUND_SEEDS before claiming a direction.\")\n\ndone = set()\nall_rows = []\nif RESUME and os.path.exists(CKPT_PATH):\n    prev = pd.read_csv(CKPT_PATH)\n    if {\"run_type\", \"species\"}.issubset(prev.columns):\n        all_rows = prev.to_dict(\"records\")\n        done = {(r[\"run_type\"], r[\"species\"]) for r in all_rows}\n        print(f\"RESUMING: {len(done)} (run_type, species) pairs already in {CKPT_PATH}\")\n\ndef _save():\n    pd.DataFrame(all_rows).to_csv(CKPT_PATH, index=False)\n\ndef _row(run_type, species, arm, n_seeds, metrics_mean, metrics_sd, **extra):\n    return dict(run_type=run_type, species=species, arm=arm, n_seeds=n_seeds,\n                auc=metrics_mean[\"auc\"], auc_sd=metrics_sd[\"auc\"],\n                f1=metrics_mean[\"f1\"], prec=metrics_mean[\"prec\"], rec=metrics_mean[\"rec\"],\n                fa_per_hr=metrics_mean[\"fa_per_hr\"], **extra)\n\nprint(f\"\\n{'='*90}\\nCELL 5  CONFOUND-CLOSING RUNS\\n{'='*90}\")\nprint(f\"species={list(CONFOUND_SPECIES)}  seeds={CONFOUND_SEEDS}  checkpoint={CKPT_PATH}\")\n\n# ======================================================================= P1: Perch alone\n# PRE-REGISTERED: Perch was pretrained as a clip-level species classifier, never as a\n# step-by-step localiser; stage-1's CNN is specifically trained per-target, few-shot, to\n# localise. Prediction: Perch-alone, as an end-to-end detector, scores LOWER than the\n# full stage-1+verify pipeline -- possibly lower than stage-1-alone too on F1/localisation\n# -- even though Perch's species-discrimination given an already-localised candidate stays\n# strong. Consistent with this project's own AP-vs-AUC finding: Perch's strength is\n# discrimination given a candidate, not candidate generation. Reported either way.\nprint(f\"\\n{'-'*90}\\nP1  PERCH ALONE AS A STANDALONE DETECTOR\\n{'-'*90}\")\nt0 = time.time(); first_timed = False\nfor sp in CONFOUND_SPECIES:\n    if (\"perch_alone\", sp) in done:\n        print(f\"  [{sp}] already checkpointed, skipping\"); continue\n    try:\n        ctx = build_target(sp, verbose=False)\n    except Exception as ex:\n        print(f\"  skip {sp}: {type(ex).__name__}: {ex}\"); continue\n    acc = {\"s1\": [], \"ver\": [], \"perch_alone\": []}\n    for sd in CONFOUND_SEEDS:\n        ts = time.time()\n        ec_f, sc_f, nc_f = eval_pass(ENC, ctx, sd), spec_pass(ENC, ctx, sd), neg_pass(ENC, ctx, sd)\n        ec_p = eval_pass(ENC, ctx, sd, detect_fn=detect_perch_only)\n        sc_p = spec_pass(ENC, ctx, sd, detect_fn=detect_perch_only)\n        nc_p = neg_pass(ENC, ctx, sd, detect_fn=detect_perch_only)\n        acc[\"s1\"].append(_arm_metrics(FUSIONS[BASE], ec_f, sc_f, nc_f, ctx))\n        acc[\"ver\"].append(_arm_metrics(FUSIONS[CAND], ec_f, sc_f, nc_f, ctx))\n        acc[\"perch_alone\"].append(_arm_metrics(FUSIONS[CAND], ec_p, sc_p, nc_p, ctx))\n        if not first_timed:\n            first_timed = True\n            per = time.time() - ts\n            print(f\"  [timing] one (species, seed) with BOTH detectors took {per:.0f}s \"\n                  f\"-> ~{per * len(CONFOUND_SPECIES) * len(CONFOUND_SEEDS) / 60:.1f} min \"\n                  f\"projected for all of P1. Widen PERCH_ALONE_STRIDE_S (CELL 2) if that's \"\n                  f\"too slow.\")\n    for arm in acc:\n        mean = {k: np.mean([d[k] for d in acc[arm]]) for k in acc[arm][0]}\n        sd_  = {k: (np.std([d[k] for d in acc[arm]], ddof=1) if len(acc[arm]) > 1 else 0.0)\n                for k in acc[arm][0]}\n        all_rows.append(_row(\"perch_alone\", sp, arm, len(CONFOUND_SEEDS), mean, sd_))\n    _save()\n    print(f\"  [{sp}] stage1-alone AUC {all_rows[-3]['auc']:.3f}  \"\n          f\"stage1+verify AUC {all_rows[-2]['auc']:.3f}  \"\n          f\"perch-alone AUC {all_rows[-1]['auc']:.3f}  [{time.time()-t0:.0f}s elapsed]\")\n    gc.collect()\n    try:\n        import torch; torch.cuda.empty_cache()\n    except Exception:\n        pass\nprint(f\"[P1 took {(time.time()-t0)/60:.1f} min]\")\n\np1 = pd.DataFrame([r for r in all_rows if r[\"run_type\"] == \"perch_alone\"])\nif len(p1):\n    g = p1.groupby(\"arm\").agg(auc=(\"auc\", \"mean\"), auc_sd=(\"auc\", \"mean\"),\n                              f1=(\"f1\", \"mean\")).reindex([\"s1\", \"ver\", \"perch_alone\"])\n    print(f\"\\n{'arm':<14}{'AUC':>8}{'F1':>8}\")\n    for arm, r in g.iterrows():\n        print(f\"{arm:<14}{r['auc']:>8.3f}{r['f1']:>8.3f}\")\n    sd_ver = float(p1[p1.arm == \"ver\"].auc_sd.mean())\n    sd_pa  = float(p1[p1.arm == \"perch_alone\"].auc_sd.mean())\n    verdict(\"Does removing stage-1 entirely change detector AUC? (verified vs perch-alone)\",\n            float(g.loc[\"ver\", \"auc\"]), float(g.loc[\"perch_alone\", \"auc\"]), sd_ver, sd_pa,\n            pos_note=\"Perch alone is BETTER than the full pipeline -- stage-1 is not \"\n                     \"adding value as a candidate proposer for these species; worth \"\n                     \"re-examining whether stage-1 is pulling its weight at all.\",\n            neg_note=\"Perch alone is WORSE than the full pipeline -- stage-1's \"\n                     \"localisation is doing real work the verifier alone cannot replace. \"\n                     \"Answers the interview question directly: the CNN's contribution is \"\n                     \"candidate generation, not species discrimination (Perch already \"\n                     \"wins that, per the AP-vs-AUC finding).\",\n            flat_note=\"Removing stage-1 doesn't move the number -- suggests Perch's own \"\n                      \"sliding-window scores are enough to localise these calls, and \"\n                      \"stage-1's contribution is smaller than assumed. Worth a wider \"\n                      \"species set before leaning on this.\")\n\n# ======================================================================= P2: species-matched negatives\n# PRE-REGISTERED: negative CHOICE is named as decisive twice already in this project's own\n# literature review (Nolasco et al. sec.4.2; Liang et al., +5.31 F1 from hard-negative\n# mining). Species-matched negatives are a strictly more informative signal than arbitrary\n# background. Prediction: stage-1-alone AUC/F1 improves with USE_SPECIES_NEG=True. No\n# strong prior on whether the gain propagates to the VERIFIED number -- the verifier\n# already handles species discrimination and may already correct for whatever stage-1\n# misses. Reported either way.\nprint(f\"\\n{'-'*90}\\nP2  STAGE-1 FINE-TUNED WITH SPECIES-MATCHED NEGATIVES\\n{'-'*90}\")\n\ndef _set_species_neg(flag):\n    g = globals()\n    g[\"USE_SPECIES_NEG\"] = flag\n    assert USE_SPECIES_NEG == flag, \"USE_SPECIES_NEG rebind failed\"\n\nt0 = time.time()\n_set_species_neg(False)   # defensive reset -- P1 never touches this flag, but don't assume\nfor sp in CONFOUND_SPECIES:\n    if (\"species_neg\", sp) in done:\n        print(f\"  [{sp}] already checkpointed, skipping\"); continue\n    try:\n        ctx = build_target(sp, verbose=False)\n    except Exception as ex:\n        print(f\"  skip {sp}: {type(ex).__name__}: {ex}\"); continue\n    if not ctx.get(\"neg_calls\"):\n        print(f\"  skip {sp}: no neg_calls in ctx (no non-decoy candidate species found)\")\n        continue\n    acc = {\"background|s1\": [], \"background|ver\": [],\n           \"species_matched|s1\": [], \"species_matched|ver\": []}\n    for use_species_neg, tag in ((False, \"background\"), (True, \"species_matched\")):\n        _set_species_neg(use_species_neg)\n        for sd in CONFOUND_SEEDS:\n            ec, sc, nc = eval_pass(ENC, ctx, sd), spec_pass(ENC, ctx, sd), neg_pass(ENC, ctx, sd)\n            acc[f\"{tag}|s1\"].append(_arm_metrics(FUSIONS[BASE], ec, sc, nc, ctx))\n            acc[f\"{tag}|ver\"].append(_arm_metrics(FUSIONS[CAND], ec, sc, nc, ctx))\n    _set_species_neg(False)   # reset before the next species / next confound section\n    for arm in acc:\n        mean = {k: np.mean([d[k] for d in acc[arm]]) for k in acc[arm][0]}\n        sd_  = {k: (np.std([d[k] for d in acc[arm]], ddof=1) if len(acc[arm]) > 1 else 0.0)\n                for k in acc[arm][0]}\n        all_rows.append(_row(\"species_neg\", sp, arm, len(CONFOUND_SEEDS), mean, sd_))\n    _save()\n    r = {rw[\"arm\"]: rw for rw in all_rows if rw[\"run_type\"] == \"species_neg\" and rw[\"species\"] == sp}\n    print(f\"  [{sp}] s1 AUC {r['background|s1']['auc']:.3f} -> {r['species_matched|s1']['auc']:.3f}   \"\n          f\"ver AUC {r['background|ver']['auc']:.3f} -> {r['species_matched|ver']['auc']:.3f}\")\n    gc.collect()\n    try:\n        import torch; torch.cuda.empty_cache()\n    except Exception:\n        pass\nprint(f\"[P2 took {(time.time()-t0)/60:.1f} min]\")\n\np2 = pd.DataFrame([r for r in all_rows if r[\"run_type\"] == \"species_neg\"])\nif len(p2):\n    for tag in (\"s1\", \"ver\"):\n        bg = p2[p2.arm == f\"background|{tag}\"]\n        sm = p2[p2.arm == f\"species_matched|{tag}\"]\n        if bg.empty or sm.empty: continue\n        verdict(f\"Does a species-matched negative pool change stage-1 fine-tuning? \"\n                f\"({'stage-1 alone' if tag == 's1' else 'verified'} AUC)\",\n                float(bg.auc.mean()), float(sm.auc.mean()),\n                float(bg.auc_sd.mean()), float(sm.auc_sd.mean()),\n                pos_note=\"Species-matched negatives HELP -- stage-1's fine-tuning negative \"\n                         \"choice matters, matching Liang et al./Nolasco et al. Worth \"\n                         \"adopting as a rung, not just a confound check.\",\n                neg_note=\"Species-matched negatives HURT -- plausible if the pool is small \"\n                         \"or dominated by acoustically-distant species, giving stage-1 \"\n                         \"less varied negative exposure than random background did.\",\n                flat_note=\"Negative source doesn't move stage-1 -- background windows were \"\n                          \"already an adequate negative signal for this task; the \"\n                          \"literature's finding may not transfer to this planting protocol.\")\n\n# ======================================================================= P3: recordist-disjoint\n# PRE-REGISTERED: if ARGUS's specificity number is genuinely about the SPECIES rather than\n# the recording SETUP, a recordist-disjoint held-out split should barely move AUC. A real\n# drop would mean part of the measured signal has been \"which microphone/recordist\", not\n# \"which bird\" -- a genuine confound this project has not yet ruled out. No strong prior\n# on magnitude either way; reported honestly regardless of outcome, per this project's\n# established practice throughout this notebook.\nprint(f\"\\n{'-'*90}\\nP3  RECORDIST-DISJOINT HELD-OUT SPLIT\\n{'-'*90}\")\nt0 = time.time()\nfor sp in CONFOUND_SPECIES:\n    if (\"recordist_disjoint\", sp) in done:\n        print(f\"  [{sp}] already checkpointed, skipping\"); continue\n    acc = {\"shared_ok|s1\": [], \"shared_ok|ver\": [],\n           \"disjoint|s1\": [], \"disjoint|ver\": []}\n    n_heldout, n_recordists = {}, {}\n    skip_species = False\n    for disjoint, tag in ((False, \"shared_ok\"), (True, \"disjoint\")):\n        try:\n            ctx = build_target(sp, recordist_disjoint=disjoint, verbose=False)\n        except (ValueError, KeyError) as ex:\n            print(f\"  [{sp}] {tag}: {type(ex).__name__}: {ex}\")\n            if disjoint:\n                print(f\"  skip {sp}: recordist-disjoint split leaves no held-out pool \"\n                      f\"(too few distinct recordists for this species) -- itself worth \"\n                      f\"logging, not a bug to route around.\")\n            skip_species = True\n            break\n        n_heldout[tag] = ctx[\"covariates\"][\"n_heldout_calls\"]\n        n_recordists[tag] = ctx[\"covariates\"][\"n_distinct_recordists\"]\n        for sd in CONFOUND_SEEDS:\n            ec, sc, nc = eval_pass(ENC, ctx, sd), spec_pass(ENC, ctx, sd), neg_pass(ENC, ctx, sd)\n            acc[f\"{tag}|s1\"].append(_arm_metrics(FUSIONS[BASE], ec, sc, nc, ctx))\n            acc[f\"{tag}|ver\"].append(_arm_metrics(FUSIONS[CAND], ec, sc, nc, ctx))\n    if skip_species:\n        continue\n    for arm in acc:\n        mean = {k: np.mean([d[k] for d in acc[arm]]) for k in acc[arm][0]}\n        sd_  = {k: (np.std([d[k] for d in acc[arm]], ddof=1) if len(acc[arm]) > 1 else 0.0)\n                for k in acc[arm][0]}\n        tag = arm.split(\"|\")[0]\n        all_rows.append(_row(\"recordist_disjoint\", sp, arm, len(CONFOUND_SEEDS), mean, sd_,\n                             n_heldout_calls=n_heldout[tag], n_distinct_recordists=n_recordists[tag]))\n    _save()\n    r = {rw[\"arm\"]: rw for rw in all_rows\n         if rw[\"run_type\"] == \"recordist_disjoint\" and rw[\"species\"] == sp}\n    print(f\"  [{sp}] shared-ok ver AUC {r['shared_ok|ver']['auc']:.3f} \"\n          f\"(n_heldout={n_heldout['shared_ok']})  ->  \"\n          f\"disjoint ver AUC {r['disjoint|ver']['auc']:.3f} \"\n          f\"(n_heldout={n_heldout['disjoint']}, distinct_recordists={n_recordists['disjoint']})\")\n    gc.collect()\n    try:\n        import torch; torch.cuda.empty_cache()\n    except Exception:\n        pass\nprint(f\"[P3 took {(time.time()-t0)/60:.1f} min]\")\n\np3 = pd.DataFrame([r for r in all_rows if r[\"run_type\"] == \"recordist_disjoint\"])\nif len(p3):\n    for tag in (\"s1\", \"ver\"):\n        so = p3[p3.arm == f\"shared_ok|{tag}\"]\n        dj = p3[p3.arm == f\"disjoint|{tag}\"]\n        if so.empty or dj.empty: continue\n        verdict(f\"Does a recordist-disjoint held-out split change AUC? \"\n                f\"({'stage-1 alone' if tag == 's1' else 'verified'})\",\n                float(so.auc.mean()), float(dj.auc.mean()),\n                float(so.auc_sd.mean()), float(dj.auc_sd.mean()),\n                pos_note=\"AUC is HIGHER once recordist overlap is removed -- unexpected; \"\n                         \"check the held-out pool sizes above before trusting this, small \"\n                         \"pools swing more.\",\n                neg_note=\"AUC FALLS once recordist overlap is removed -- part of the \"\n                         \"measured signal was 'recognise this recordist's setup', not \"\n                         \"purely the species. A real, reportable limitation -- state the \"\n                         \"size of the drop, don't just flag its direction.\",\n                flat_note=\"AUC does not depend on recordist overlap -- the specificity \"\n                          \"result is not an artefact of recording-setup leakage. Rules \"\n                          \"out the 'did it learn the microphone, not the bird' concern \"\n                          \"for these species.\")\n\nprint(f\"\\n{'='*90}\\nALL CONFOUND-CLOSING RUNS DONE -- {len(all_rows)} rows -> {CKPT_PATH}\")\nprint(f\"Log the three verdicts above in the bound logbook the same day.\\n{'='*90}\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"23a79fe9","cell_type":"markdown","source":"## After running\n\n- `argus_multispecies_results.csv`, `argus_ablation_ladder.csv`,\n  `argus_sweep_shot_strategy.csv`, `argus_sweep_shot_count.csv`,\n  `argus_sweep_snr.csv`, `argus_sweep_decoy_difficulty.csv`, `argus_confounds.csv` are\n  written to `/kaggle/working/`. Download them (or commit the notebook) before the\n  session ends.\n- Log the headline numbers — per-species AUC before/after, the winning ladder\n  config, which Perch version was used, and Cell 5's three confound verdicts — in the\n  bound logbook **the same day**.\n- If this was a `SMOKE_TEST` run: record hours-per-species-per-seed, then come back,\n  set `SMOKE_TEST = False`, and size `SEEDS` / `N_TARGETS` to the real budget before\n  the next run.\n","metadata":{}}]}