{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13"},"accelerator":"GPU"},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"8bba2726-f9c4-499e-bf64-bdbac65073be","cell_type":"markdown","source":"# NB-05b — 1 h Trajectory Attacks (locked configuration)\n\n**One knob. Everything else is hardcoded and asserted.** This notebook exists because the\ngeneral-purpose NB-05 has eleven settings, and on 2026-09-26 four runs (~1.6 GPU-h) were lost\nto one of them being wrong in a way nothing checked.\n\n---\n\n## What this notebook achieves\n\nIt measures targeted attack success at a **fixed** perturbation budget for the four checkpoints\nof the 1 h trajectory, so those four conditions can be **pooled with the eight 25h/100h\ntrajectory conditions already measured**. Pooling is the entire point, and pooling is only\nvalid if every setting below matches those eight runs exactly.\n\n| Why it matters | |\n|---|---|\n| **Step-matched contrasts** | 1h step-N vs 100h step-N holds optimiser steps — and therefore training stage — constant **by construction**, so earliness cannot explain a difference. Exposure ratios: 40x (step 1000), 79x (2000), 100x (3500), 100x (final). |\n| **Exposure range** | The 8-condition fit spanned only 24.6–98 h (4x) and gave `log10(exposure)` p = 0.117. Twelve conditions span 0.98–98 h (**100x**), which is where RQ1's effects live. |\n| **Third earliness curve** | Already answered on the wrong eval set: the 1 h rung shows **no** trajectory (ASR range 13.2 pp, non-monotonic, all p > 0.34) because one epoch is 25 steps, so step-1000 is the 40th epoch. This run confirms it on the RQ1 utterance set. |\n\n**What it does NOT achieve:** a new iso-WER pair. The 1 h trajectory spans attack WER ~0.80–0.647\nand the 100 h trajectory 0.667–0.250; they overlap only where pair 1 already sits. Clean WER is\ndeliberately *unmatched* in the step-matched contrasts and enters NB-09's ANCOVA as a covariate.\nDo not go looking for a fifth pair.\n\n---\n\n## Kaggle settings\n\n| Setting | Value |\n|---|---|\n| Accelerator | **GPU T4 x2** |\n| Internet | **On** |\n| Persistence | Variables and Files |\n\n### Inputs — exactly these three\n\n| # | Dataset | Holds |\n|---|---|---|\n| 1 | `bengaliai-speech` | the `.mp3` audio. The split CSVs carry only *relative* paths. |\n| 2 | `bn-asr-splits` | `attack_eval_25.csv`, `test.csv`, `vocab.json` |\n| 3 | `ckpt-xlsr-1h-traj` | the four 1 h checkpoints |\n\n### Do NOT attach `bn-attack-targets`\n\nThat dataset is for RQ2. Its utterance set (`attack_eval_t1.csv`, 38 utts) is **disjoint** from\nthis one (`attack_eval_25.csv`, 25 utts) — one has zero standalone-\"না\" utterances, the other is\nnothing but. Attaching it is what broke the first attempt. This notebook now **ignores it even\nif attached** and the preflight gate says so, but leave it off anyway.\n\n---\n\n## The four runs\n\nChange **`CHECKPOINT` only**. Nothing else is editable.\n\n| # | `CHECKPOINT` | Publishes as |\n|---|---|---|\n| 1 | `\"milestone-1000\"` | `eps50b-rung1-ms1000-fixedeps-t5` |\n| 2 | `\"milestone-2000\"` | `eps50b-rung1-ms2000-fixedeps-t5` |\n| 3 | `\"milestone-3500\"` | `eps50b-rung1-ms3500-fixedeps-t5` |\n| 4 | `\"final\"` | `eps50b-rung1-final-fixedeps-t5` |\n\nThe `eps50b-` prefix keeps these separate from the four superseded 38-utt datasets, which are\nkept for the record. ~25 min each, ~1.6 GPU-h total. Two sessions in parallel is fine.\n\n---\n\n## Verify in the log before trusting a run\n\nThe preflight gate (section 4b) checks all of this and **aborts** if any of it is wrong, but\nread these lines anyway:\n\n1. `eval set [t5_frozen]: .../attack_eval_25.csv` — **not** `attack_eval_t1.csv`\n2. `T5: 25 attackable utterances` — **25, not 38**\n3. `targets: AUTO-GENERATED (RandomState(1337))` — not the NB-03 corpus\n4. `checkpoint: .../ckpt-xlsr-1h-traj/milestones/step-...`\n5. `PREFLIGHT PASSED — comparable to the 8 baseline conditions`\n6. `Converged. eps50 is stable at this budget`\n","metadata":{}},{"id":"3cd6ba70-9097-4de2-8f1b-ae73df09d492","cell_type":"markdown","source":"## 1 · Configuration — one knob, the rest locked","metadata":{}},{"id":"6bf08ad4-b906-4b7a-b416-382dd2a1faac","cell_type":"code","source":"# ============================================================================\n#  THE ONLY SETTING YOU CHANGE\n# ============================================================================\nCHECKPOINT = \"milestone-3500\"    # <<<< \"milestone-1000\" | \"milestone-2000\"\n                                 #      \"milestone-3500\" | \"final\"\n\n# ============================================================================\n#  LOCKED — do not edit. These are the eight baseline trajectory conditions'\n#  settings. Changing any of them makes this run incomparable to those eight,\n#  which destroys the only reason the notebook exists. The preflight gate in\n#  section 4b re-checks every one of them against the measured state.\n# ============================================================================\nRUNG         = 1                    # the 1 h rung\nEDIT_TYPE    = \"T5\"                 # full retarget, as the 8 baseline runs used\nEVAL_SET     = \"t5_frozen\"          # attack_eval_25.csv, 25 utts -- THE RQ1 SPINE.\n                                    # \"nb03\" would silently give the 38-utt RQ2 set,\n                                    # which is DISJOINT from it. That was the\n                                    # 2026-09-26 defect.\nFIXED_EPS    = 2e-3                 # same budget as the 8 baseline runs\nCKPT_DATASET = \"ckpt-xlsr-1h-traj\"  # mandatory: ckpt-xlsr-1h has no milestones/\nN_EVAL       = 25                   # selects attack_eval_25.csv\n\n_LOCKED = dict(RUNG=1, EDIT_TYPE=\"T5\", EVAL_SET=\"t5_frozen\", FIXED_EPS=2e-3,\n               CKPT_DATASET=\"ckpt-xlsr-1h-traj\", N_EVAL=25)\n_actual = dict(RUNG=RUNG, EDIT_TYPE=EDIT_TYPE, EVAL_SET=EVAL_SET,\n               FIXED_EPS=FIXED_EPS, CKPT_DATASET=CKPT_DATASET, N_EVAL=N_EVAL)\nassert _actual == _LOCKED, (\n    \"a LOCKED setting was edited, so this run would not be comparable to the \"\n    f\"8 baseline trajectory conditions.\\n  expected {_LOCKED}\\n  got      {_actual}\")\nassert CHECKPOINT in (\"milestone-1000\", \"milestone-2000\", \"milestone-3500\", \"final\"), \\\n    f\"CHECKPOINT={CHECKPOINT!r} is not one of the four 1 h trajectory checkpoints\"\nprint(f\"NB-05b  1 h trajectory  |  CHECKPOINT = {CHECKPOINT}\")\nprint(\"locked config OK\\n\")\n\n# --- unchanged operational settings ---------------------------------------\nBATCH_SIZE  = 8          # per-batch utterances; drop to 4 if you hit OOM\nUSE_AMP     = True       # settled by NB-04b run 2: AMP reproduced fp32 eps50 exactly\nSESSION_HOURS = 8.0      # Kaggle kills a committed GPU session at 12 h\nVALIDATE_BATCHING = False\n\nEPS_LO, EPS_HI = (1e-5, 2e-3) if EDIT_TYPE == \"T1\" else (1e-4, 2e-3)\nBISECT_ROUNDS  = 6 if EDIT_TYPE == \"T1\" else 5\nif FIXED_EPS is not None:\n    EPS_HI = float(FIXED_EPS)\n    EPS_LO = min(EPS_LO, EPS_HI)\n    BISECT_ROUNDS = 0\nPROBE_STEPS    = 400     # steps per bisection probe (matches NB-04b)\nCONFIRM_STEPS  = 800     # steps for the final confirmation at the chosen eps\nSTALL_PATIENCE = 100     # abort a probe after this many steps with no loss improvement\nCHECK_EVERY    = 50\nLR             = 1e-4\n\n# Costs roughly THREE extra eps50 searches (~1-2 GPU-h) because it runs the\n# same two utterances batched and then one at a time. Run it ONCE, for\n# (RUNG=1, EDIT_TYPE=\"T5\"), record the agreement in RUNLOG.md, then set it\n# False for the remaining seven runs. It is a one-off correctness check on the\n# batching, not something each run needs to repeat.\nVALIDATE_BATCHING = False   # <<<< set True for ONE run only, then back to False.\n                            # NOTE: RUNLOG has no row recording the batching\n                            # agreement figure. It was almost certainly measured\n                            # on the first run (RUNG=1, T5 -- the old defaults),\n                            # but check that run's log and add the number to\n                            # RUNLOG before the thesis cites batched eps50.","metadata":{},"outputs":[],"execution_count":null},{"id":"11c483d1-656a-410d-a862-f8387a08c70b","cell_type":"markdown","source":"## 2 · Boilerplate","metadata":{}},{"id":"e0c077f0-58ec-44db-a403-7749a8fee3a2","cell_type":"code","source":"import os\nos.environ[\"CUDA_VISIBLE_DEVICES\"] = \"0\"   # attacks optimise the INPUT: never DataParallel\nimport sys, json, time, random, re, glob, unicodedata\nimport numpy as np, pandas as pd, torch\nimport torch.nn.functional as F\ntorch.backends.cudnn.benchmark = True\n\nSEED = 1337\nrandom.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)\nWORK, TEMP, DEVICE = \"/kaggle/working\", \"/kaggle/temp\", \"cuda\"\nassert torch.cuda.is_available(), \"Settings -> Accelerator -> GPU T4 x2\"\n\nPROV = {\"notebook\": \"NB-05b-traj-1h-locked\", \"rung\": RUNG, \"edit_type\": EDIT_TYPE,\n        \"n_eval\": N_EVAL, \"use_amp\": USE_AMP, \"seed\": SEED,\n        \"bracket\": [EPS_LO, EPS_HI], \"bisect_rounds\": BISECT_ROUNDS,\n        \"gpu\": torch.cuda.get_device_name(0), \"torch\": torch.__version__,\n        \"timestamp\": time.strftime(\"%Y-%m-%d %H:%M:%S\")}\nprint(json.dumps(PROV, indent=2))","metadata":{},"outputs":[],"execution_count":null},{"id":"5f4acc6e-6822-4067-b5a4-70f6b7004c28","cell_type":"code","source":"# See NB-01 for the full explanation: pinning an old `accelerate`/`datasets`\n# silently downgrades numpy mid-session and breaks pandas much later with\n# \"No module named 'numpy.rec'\". Pin numpy to what is already loaded.\n# `datasets` is not imported by this notebook, so it is not installed.\nimport numpy, subprocess, sys\nNP_BEFORE = numpy.__version__\n\n!pip install -q transformers==4.44.2 librosa soundfile jiwer \"numpy=={NP_BEFORE}\" 2>&1 | tail -2\n\nNP_DISK = subprocess.run([sys.executable, \"-c\", \"import numpy; print(numpy.__version__)\"],\n                         capture_output=True, text=True).stdout.strip()\nassert NP_BEFORE == NP_DISK, f\"numpy moved {NP_BEFORE} -> {NP_DISK}; relax the pin that forced it.\"\nprint(f\"numpy {NP_BEFORE} stable\")\n\nfrom transformers import Wav2Vec2ForCTC, Wav2Vec2Processor\nimport transformers, librosa\nprint(\"transformers\", transformers.__version__)\nPROV[\"transformers\"] = transformers.__version__\nPROV[\"numpy\"] = NP_BEFORE\n","metadata":{},"outputs":[],"execution_count":null},{"id":"580a920e-d5f9-4101-a9e9-2e9cd2abe5f0","cell_type":"markdown","source":"## 3 · Load the ladder checkpoint and the frozen attack set","metadata":{}},{"id":"09d6b79f-d877-4ae5-9382-61d72b545a50","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Kaggle does not use one mount layout. Depending on the kernel you get\n#     /kaggle/input/<name>                (older kernels)\n#     /kaggle/input/datasets/<name>       (your own datasets, current)\n#     /kaggle/input/competitions/<name>   (competition data)\n# A flat glob only sees the wrapper dirs -- NB-01 failed with exactly this,\n# reporting: Visible inputs: ['competitions', 'datasets']\n# ---------------------------------------------------------------------------\ninputs = sorted({p for pat in (\"/kaggle/input/*\", \"/kaggle/input/*/*\",\n                               \"/kaggle/input/*/*/*\")\n                 for p in glob.glob(pat) if os.path.isdir(p)})\n\n# Match the rung EXACTLY. The old test was `f\"{RUNG}h\" in p`, and \"5h\" is a\n# substring of \"ckpt-xlsr-25h\" -- with both datasets attached, RUNG=5 could\n# silently attack the 25 h model and corrupt the headline figure.\n_DS = CKPT_DATASET or f\"ckpt-xlsr-{RUNG}h\"\ncands = [p for p in inputs if re.search(rf\"{re.escape(_DS)}$\", os.path.basename(p))]\nassert len(cands) == 1, (\n    f\"Need exactly one input matching '{_DS}', found {len(cands)}: \"\n    f\"{[os.path.basename(c) for c in cands]}\\n\"\n    f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\nCKPT = cands[0]\nif CHECKPOINT == \"final\":\n    MODEL_DIR = CKPT if os.path.exists(f\"{CKPT}/config.json\") else f\"{CKPT}/final\"\nelse:\n    MODEL_DIR = f\"{CKPT}/milestones/step-{CHECKPOINT.split('-')[1]}\"\n    assert os.path.isdir(MODEL_DIR), (\n        f\"{MODEL_DIR} missing. Milestones live at steps 1000/2000/3500. They exist \"\n        f\"for rungs >= 25h in ckpt-xlsr-{{25,100}}h, and for the 1h rung ONLY in \"\n        f\"ckpt-xlsr-1h-traj (the 2026-09-26 TRAJ_RUN retrain) -- the original \"\n        f\"ckpt-xlsr-1h has no milestones/ dir. If you want 1h milestones, set \"\n        f\"CKPT_DATASET = 'ckpt-xlsr-1h-traj'.\")\nassert os.path.exists(f\"{MODEL_DIR}/config.json\"), \\\n    f\"no config.json under {MODEL_DIR} -- did NB-01 finish and publish this rung?\"\nprint(\"checkpoint:\", MODEL_DIR)\n\nprocessor = Wav2Vec2Processor.from_pretrained(MODEL_DIR)\nmodel = Wav2Vec2ForCTC.from_pretrained(MODEL_DIR).float().to(DEVICE).eval()\nfor p in model.parameters(): p.requires_grad_(False)\nmodel.config.apply_spec_augment = False      # no masking noise during attacks\n\nDO_NORMALIZE = bool(getattr(processor.feature_extractor, \"do_normalize\", True))\nBLANK_ID = model.config.pad_token_id\nprint(f\"do_normalize {DO_NORMALIZE} | blank {BLANK_ID} | vocab {model.config.vocab_size}\")\n\nSPLITS = next((p for p in inputs if \"splits\" in p.lower()), None)\nassert SPLITS, (f\"Attach bn-asr-splits. Visible: {[os.path.basename(p) for p in inputs]}\")\n# Which utterances to attack. The NB-03 set is used for BOTH edit types when it\n# is attached: T1 vs T5 is only a paired comparison if both run on the same audio.\n# attack_eval_25.csv has 0 standalone-\"না\" utterances (verified 2026-09-25), so T1\n# can never use it.\nassert EVAL_SET in (\"t5_frozen\", \"nb03\"), f\"EVAL_SET={EVAL_SET!r} is not valid\"\n_frozen = f\"{SPLITS}/attack_eval_{N_EVAL}.csv\"\n_nb03 = next((f\"{p}/attack_eval_t1.csv\" for p in inputs\n              if os.path.exists(f\"{p}/attack_eval_t1.csv\")), None)\nif EVAL_SET == \"nb03\":\n    assert _nb03, (\n        \"EVAL_SET='nb03' but attack_eval_t1.csv is not attached. Attach \"\n        \"bn-attack-targets (run NB-03 first).\\n\"\n        f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n    EV_CSV = _nb03\nelse:\n    assert EDIT_TYPE != \"T1\", (\n        \"EVAL_SET='t5_frozen' cannot run T1: attack_eval_25.csv has zero \"\n        \"standalone-'na' utterances (verified 2026-09-25). Use EVAL_SET='nb03'.\")\n    assert os.path.exists(_frozen), f\"{_frozen} missing -- attach bn-asr-splits\"\n    EV_CSV = _frozen\n    if _nb03:\n        print(\"note: bn-attack-targets is attached but EVAL_SET='t5_frozen', so it is \"\n              \"IGNORED -- both the utterance set and the targets come from \"\n              \"attack_eval_25.csv. Deliberate: attaching a dataset must never \"\n              \"change which utterances get attacked.\")\nprint(f\"eval set [{EVAL_SET}]: {EV_CSV}\")\nev = pd.read_csv(EV_CSV)\nPROV[\"attack_eval_csv\"] = os.path.basename(EV_CSV)\nPROV[\"fixed_eps\"] = FIXED_EPS\nPROV[\"ckpt_dataset\"] = _DS\nPROV[\"bisect_rounds\"] = BISECT_ROUNDS\n\n# ---------------------------------------------------------------------------\n# Re-resolve the audio root, same hazard as NB-01: NB-00 baked in the mount\n# point from ITS session. Here the whole attack set is only 25 files, so check\n# every one of them rather than sampling.\n# ---------------------------------------------------------------------------\ndef resolve_audio(df):\n    p0 = str(df[\"path\"].iloc[0])\n    if os.path.exists(p0):\n        return df, os.path.dirname(p0)\n    stem, ext = os.path.basename(os.path.dirname(p0)), os.path.splitext(p0)[1]\n    cands = [d for pat in (f\"/kaggle/input/*/{stem}\",\n                           f\"/kaggle/input/*/*/{stem}\",\n                           f\"/kaggle/input/*/*/*/{stem}\")\n             for d in glob.glob(pat) if os.path.isdir(d)]\n    assert cands, (f\"Audio directory '{stem}' not found under /kaggle/input.\\n\"\n                   \"Attach: + Add Input -> Competitions -> 'bengaliai speech'.\\n\"\n                   f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n    df = df.copy()\n    df[\"path\"] = cands[0] + \"/\" + df[\"utt_id\"].astype(str) + ext\n    return df, cands[0]\n\nev, AUDIO_ROOT = resolve_audio(ev)\nabsent = [p for p in ev[\"path\"] if not os.path.exists(p)]\nassert not absent, f\"{len(absent)}/{len(ev)} attack audio files missing, e.g.\\n  {absent[0]}\"\nPROV[\"model_dir\"] = MODEL_DIR; PROV[\"audio_root\"] = AUDIO_ROOT\nprint(f\"audio root: {AUDIO_ROOT}\")\nprint(f\"attack set: {len(ev)} utterances, all present\")","metadata":{},"outputs":[],"execution_count":null},{"id":"93ff2db4-bbf7-43c2-ad72-beb152036586","cell_type":"markdown","source":"## 4 · Build the targets\n\nIf NB-03's hand-built corpus is attached, use it. Otherwise generate automatically:\n\n* **T1** — delete `না` from the reference transcript. Post-verbal, short, unstressed; one\n  syllable inverts the meaning.\n* **T5** — a *different real Bengali sentence* drawn from the corpus with a **matched token\n  count**. This matters: NB-04's T5 target was 9 tokens against a ~55-character source, and\n  forcing a much shorter output is plausibly easier, which would flatter T5 and bias against\n  the T1 hypothesis. Matching length removes that.","metadata":{}},{"id":"4baf9c69-ee26-4ec8-ac7d-0dfe92f8e06b","cell_type":"code","source":"ZW = dict.fromkeys(map(ord, \"‌‍\"), None)\ndef norm_bn(s):\n    return re.sub(r\"\\s+\", \" \", unicodedata.normalize(\"NFC\", str(s)).translate(ZW)).strip()\ndef ntok(s): return len(processor.tokenizer(norm_bn(s)).input_ids)\n\n# Use the resolved `inputs` list from the cell above, NOT a flat glob. Your own\n# datasets mount at /kaggle/input/datasets/<user>/<name>, three levels deep, so\n# glob(\"/kaggle/input/*\") sees only the wrapper dirs. That is the same failure\n# NB-01 hit, and on 2026-09-25 it silently hid an attached bn-attack-targets:\n# T5 then auto-generated targets on the wrong utterance set and still \"succeeded\".\n# The TARGET source is tied to EVAL_SET so the two can never diverge. Under\n# \"t5_frozen\" the NB-03 corpus is not consulted even when attached: its utt_ids\n# are disjoint from attack_eval_25.csv, so a merge would inner-join to zero rows.\nTGT_CSV = None\nif EVAL_SET == \"nb03\":\n    TGT_CSV = next((f\"{p}/bn_attack_targets.csv\" for p in inputs\n                    if os.path.exists(f\"{p}/bn_attack_targets.csv\")), None)\n    assert TGT_CSV, (\n        \"EVAL_SET='nb03' but bn_attack_targets.csv not found. Attach \"\n        \"bn-attack-targets (run NB-03 first).\\n\"\n        f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n\nif TGT_CSV:\n    tg = pd.read_csv(TGT_CSV)\n    tg = tg[tg.edit_type == EDIT_TYPE]\n    ev = ev.merge(tg[[\"source_utterance_id\", \"target_text\"]],\n                  left_on=\"utt_id\", right_on=\"source_utterance_id\", how=\"inner\")\n    print(f\"using NB-03 hand-built corpus: {len(ev)} pairs\")\nelse:\n    print(\"targets: AUTO-GENERATED from test.csv, RandomState(1337) -- CORRECT for\")\n    print(\"         NB-05b. The eight 25h/100h baseline trajectory conditions were\")\n    print(\"         measured this way, so matching them is the whole point. Do NOT\")\n    print(\"         substitute the NB-03 hand corpus here: its utterances are\")\n    print(\"         DISJOINT from attack_eval_25.csv and the preflight gate will\")\n    print(\"         refuse the run.\\n\")\n    pool = pd.read_csv(f\"{SPLITS}/test.csv\")\n    pool[\"ntok\"] = pool[\"text\"].map(ntok)\n    ev[\"src_ntok\"] = ev[\"text\"].map(ntok)\n\n    rows = []\n    rng = np.random.RandomState(SEED)\n    for _, r in ev.iterrows():\n        srct = norm_bn(r[\"text\"])\n        if EDIT_TYPE == \"T1\":\n            # STANDALONE particle only. A substring test is wrong: \"না\" occurs\n            # inside many common Bengali words -- নাম (name), নারী (woman),\n            # জানা (to know) -- and str.replace() would carve syllables out of\n            # them, producing a target that is not a negation inversion at all.\n            # Remove the LAST standalone \"না\", which is the post-verbal negation.\n            toks = srct.split()\n            if \"না\" not in toks: continue\n            j = len(toks) - 1 - toks[::-1].index(\"না\")\n            tgt = norm_bn(\" \".join(toks[:j] + toks[j+1:]))\n            if tgt == srct or not tgt: continue\n        else:\n            cand = pool[(pool.utt_id != r[\"utt_id\"]) &\n                        (pool.ntok.between(r[\"src_ntok\"] - 2, r[\"src_ntok\"] + 2))]\n            if len(cand) == 0: continue\n            tgt = norm_bn(cand.sample(1, random_state=int(rng.randint(1e6))).iloc[0][\"text\"])\n        rows.append({**r.to_dict(), \"target_text\": tgt})\n    ev = pd.DataFrame(rows)\n    if EDIT_TYPE == \"T1\":\n        # NB-00 froze the attack set using a SUBSTRING test for \"না\", so the\n        # advertised \"15 with না\" overcounts utterances that actually carry the\n        # standalone negation particle. Find out before spending GPU.\n        print(f\"\\nT1: {len(ev)} of {N_EVAL} utterances carry a standalone না\")\n        assert len(ev) >= 10, (\n            f\"only {len(ev)} T1-attackable utterances -- too few for RQ2. \"\n            f\"Either build the NB-03 corpus, or freeze an additional T1-specific \"\n            f\"attack subset from the test split (without touching bn-asr-splits).\")\n\nev[\"target_text\"] = ev[\"target_text\"].map(norm_bn)\nev[\"target_tokens\"] = ev[\"target_text\"].map(ntok)\nev[\"source_tokens\"] = ev[\"text\"].map(ntok)\nev = ev.reset_index(drop=True)\nPROV[\"n_attackable\"] = int(len(ev))\nPROV[\"n_eval_frozen\"] = int(N_EVAL)\nPROV[\"mean_target_tokens\"] = float(ev.target_tokens.mean())\nPROV[\"mean_source_tokens\"] = float(ev.source_tokens.mean())\nprint(f\"{EDIT_TYPE}: {len(ev)} attackable utterances\")\nprint(f\"target tokens  mean {ev.target_tokens.mean():.1f}  vs source {ev.source_tokens.mean():.1f}\")\nprint(\"  (these should be close for T5; T1 is by construction slightly shorter)\")\nev[[\"utt_id\", \"source_tokens\", \"target_tokens\", \"text\", \"target_text\"]].head()","metadata":{},"outputs":[],"execution_count":null},{"id":"b96e56f3-a3ea-4f9a-972b-260fbc0ead6b","cell_type":"markdown","source":"## 4b · Preflight gate — would this run be comparable to the baseline eight?","metadata":{}},{"id":"392bfb48-2860-4638-9fc1-2959d480b894","cell_type":"code","source":"# ============================================================================\n#  PREFLIGHT GATE\n#  Every assertion here corresponds to something that has actually gone wrong.\n#  A consequence-based gate: it checks the MEASURED state (which CSV was really\n#  opened, how many utterances really loaded, where the targets really came\n#  from), never just the declared config, because the 2026-09-26 defect was a\n#  declared config that was fine and a resolved state that was not.\n# ============================================================================\n_fail = []\n\n# 1. THE defect of 2026-09-26: the eval set flipped to the 38-utt RQ2 corpus\n#    because bn-attack-targets was attached. Check the file actually opened.\nif os.path.basename(EV_CSV) != \"attack_eval_25.csv\":\n    _fail.append(f\"eval set is {os.path.basename(EV_CSV)!r}, must be \"\n                 f\"'attack_eval_25.csv'. The 38-utt NB-03 set is DISJOINT from it \"\n                 f\"(zero standalone-na utterances vs nothing but), so results would \"\n                 f\"not be comparable to the 8 baseline conditions. Detach \"\n                 f\"bn-attack-targets.\")\n\n# 2. Utterance count, measured after the merge/generation that builds `ev`.\nif len(ev) != 25:\n    _fail.append(f\"{len(ev)} utterances loaded, expected exactly 25. 38 means the \"\n                 f\"NB-03 corpus was used.\")\n\n# 3. Targets must be auto-generated, deterministically, from test.csv -- the\n#    same source the 8 baseline runs used. TGT_CSV set means the hand corpus won.\nif TGT_CSV is not None:\n    _fail.append(f\"targets came from {os.path.basename(TGT_CSV)}; this run must use \"\n                 f\"auto-generated targets (RandomState(1337)) to match the baseline.\")\n\n# 4. Right weights.\nif \"ckpt-xlsr-1h-traj\" not in MODEL_DIR:\n    _fail.append(f\"MODEL_DIR={MODEL_DIR!r} is not in ckpt-xlsr-1h-traj. The original \"\n                 f\"ckpt-xlsr-1h has no milestones/ directory.\")\n_want = \"final\" if CHECKPOINT == \"final\" else f\"step-{CHECKPOINT.split('-')[1]}\"\nif not MODEL_DIR.rstrip(\"/\").endswith(_want):\n    _fail.append(f\"MODEL_DIR={MODEL_DIR!r} does not end in {_want!r}\")\n\n# 5. FIXED_EPS mode must really be active: bisection off, ceiling at the budget.\nif BISECT_ROUNDS != 0:\n    _fail.append(f\"BISECT_ROUNDS={BISECT_ROUNDS}, expected 0 under FIXED_EPS\")\nif abs(EPS_HI - 2e-3) > 1e-12:\n    _fail.append(f\"EPS_HI={EPS_HI:.3e}, expected 2.000e-03\")\n\nassert not _fail, \"PREFLIGHT FAILED:\\n  - \" + \"\\n  - \".join(_fail)\n\n_EXPOSURE_100H = {\"milestone-1000\": 38.8, \"milestone-2000\": 77.7,\n                  \"milestone-3500\": 98.0, \"final\": 98.0}\n_ratio = _EXPOSURE_100H[CHECKPOINT] / 0.98\nprint(\"PREFLIGHT PASSED -- comparable to the 8 baseline conditions\")\nprint(f\"  eval set        : {os.path.basename(EV_CSV)}  ({len(ev)} utterances)\")\nprint(f\"  targets         : AUTO-GENERATED (RandomState(1337)) from test.csv\")\nprint(f\"  checkpoint      : {CHECKPOINT}\")\nprint(f\"  budget          : FIXED_EPS {EPS_HI:.1e}, bisection off\")\nprint()\nprint(\"  this run contributes one condition to the 12-condition pool:\")\nprint(f\"    exposure           0.98 h (flat: one epoch at the 1 h rung is 25 steps,\")\nprint(f\"                       so every checkpoint has seen all 809 clips)\")\nprint(f\"    step-matched vs    100h {CHECKPOINT} at {_EXPOSURE_100H[CHECKPOINT]} h \"\n      f\"-> {_ratio:.0f}x exposure, identical optimiser steps\")\nprint(f\"    reported metric    ASR@{EPS_HI:.0e} over all {len(ev)} utterances \"\n      f\"(censoring-free)\")\nprint()\nprint(\"  NOT produced by this run: a new iso-WER pair. Clean WER is deliberately\")\nprint(\"  unmatched here and enters NB-09's ANCOVA as a covariate instead.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"5601bdd2-8ab2-4246-afd7-a459551f652c","cell_type":"markdown","source":"## 5 · Batched differentiable front-end and loss","metadata":{}},{"id":"c5c32abb-3c84-4c1f-a762-c26ce8d8c809","cell_type":"code","source":"SR = 16000\n\ndef load_batch(rows):\n    \"\"\"Length-bucketed batch. Returns padded waveforms + a validity mask.\"\"\"\n    wavs = []\n    for _, r in rows.iterrows():\n        a, _ = librosa.load(r[\"path\"], sr=SR, mono=True)\n        a = np.asarray(a, np.float32)\n        pk = np.abs(a).max()\n        wavs.append(a / max(pk, 1.0))\n    T = max(len(w) for w in wavs)\n    x = torch.zeros(len(wavs), T, dtype=torch.float32)\n    m = torch.zeros(len(wavs), T, dtype=torch.float32)\n    for i, w in enumerate(wavs):\n        x[i, :len(w)] = torch.from_numpy(w); m[i, :len(w)] = 1.0\n    return x.to(DEVICE), m.to(DEVICE)\n\n\ndef frontend_b(x, mask):\n    \"\"\"Masked zero-mean/unit-variance. Computing stats over the PADDING would\n    corrupt short utterances in a mixed-length batch.\"\"\"\n    if DO_NORMALIZE:\n        n = mask.sum(-1, keepdim=True).clamp(min=1)\n        mu = (x * mask).sum(-1, keepdim=True) / n\n        var = (((x - mu) * mask) ** 2).sum(-1, keepdim=True) / n\n        x = (x - mu) / torch.sqrt(var + 1e-7)\n        x = x * mask\n    return x\n\n\ndef logits_b(x, mask):\n    return model(input_values=frontend_b(x, mask),\n                 attention_mask=mask.long()).logits\n\n\ndef ctc_per_sample(x, mask, tgt_pad, tgt_len):\n    \"\"\"Per-sample CTC loss -> lets us mask out finished samples.\"\"\"\n    lg = logits_b(x, mask)\n    logp = F.log_softmax(lg.float(), -1).transpose(0, 1)          # (L,B,V)\n    in_len = model._get_feat_extract_output_lengths(mask.sum(-1).long()).long()\n    in_len = in_len.clamp(max=logp.shape[0])\n    return F.ctc_loss(logp, tgt_pad, in_len, tgt_len,\n                      blank=BLANK_ID, zero_infinity=True, reduction=\"none\")\n\n\ndef decode_b(x, mask):\n    with torch.no_grad():\n        ids = torch.argmax(logits_b(x, mask), -1)\n    return [norm_bn(s) for s in processor.batch_decode(ids)]\n\n\ndef int16_rt(x):\n    return torch.clamp((x * 32767.0).round(), -32768.0, 32767.0) / 32767.0\n\ndef snr_db_b(clean, adv, mask):\n    n = (adv - clean) * mask\n    s = clean * mask\n    return (10 * torch.log10(s.pow(2).sum(-1) / (n.pow(2).sum(-1) + 1e-12))).cpu().numpy()","metadata":{},"outputs":[],"execution_count":null},{"id":"d4cc1358-2bfe-4ffb-a1ae-c237a096dccb","cell_type":"markdown","source":"## 6 · Batched PGD with per-sample budgets\n\nEvery utterance in the batch carries its **own** eps, its own success flag and its own\nbest-so-far delta. Finished samples are masked out of the loss so the optimiser stops spending\ncapacity on them, and their delta is frozen at the moment of success.","metadata":{}},{"id":"afaed727-6740-43bc-944e-2459078f1753","cell_type":"code","source":"def pgd_batch(x, mask, tgt_pad, tgt_len, targets, eps_vec,\n              steps, delta_init=None, lr=LR):\n    \"\"\"eps_vec: (B,) per-sample L-inf budgets.\n       Returns (success bool array, frozen deltas, steps_used).\"\"\"\n    B = x.shape[0]\n    eps_col = eps_vec.view(-1, 1)\n    delta = (torch.zeros_like(x) if delta_init is None\n             else torch.max(torch.min(delta_init.clone(), eps_col), -eps_col))\n    delta.requires_grad_(True)\n    opt = torch.optim.Adam([delta], lr=lr)\n    scaler = torch.amp.GradScaler(\"cuda\", enabled=USE_AMP)\n\n    active = torch.ones(B, dtype=torch.bool, device=DEVICE)\n    frozen = torch.zeros_like(x)\n    success = np.zeros(B, dtype=bool)\n    best_loss = torch.full((B,), float(\"inf\"), device=DEVICE)\n    stall = torch.zeros(B, dtype=torch.long, device=DEVICE)\n\n    for step in range(steps):\n        if not active.any():\n            break\n        x_adv = torch.clamp(x + delta, -1, 1) * mask\n        with torch.amp.autocast(\"cuda\", enabled=USE_AMP):\n            per = ctc_per_sample(x_adv, mask, tgt_pad, tgt_len)\n        loss = (per * active.float()).sum() / active.float().sum().clamp(min=1)\n\n        opt.zero_grad(); scaler.scale(loss).backward()\n        scaler.unscale_(opt)\n        if delta.grad is not None:\n            delta.grad[~active] = 0          # finished samples stop moving\n        scaler.step(opt); scaler.update()\n        with torch.no_grad():\n            delta.data = torch.max(torch.min(delta.data, eps_col), -eps_col)\n\n        with torch.no_grad():                # stall detection\n            improved = per < best_loss - 1e-3\n            best_loss = torch.minimum(best_loss, per.detach())\n            stall = torch.where(improved, torch.zeros_like(stall), stall + 1)\n            give_up = active & (stall > STALL_PATIENCE)\n            active = active & ~give_up\n\n        if (step + 1) % CHECK_EVERY == 0:\n            hyp = decode_b(int16_rt(torch.clamp(x + delta, -1, 1)) * mask, mask)\n            with torch.no_grad():\n                for i in range(B):\n                    if active[i] and hyp[i] == targets[i]:\n                        frozen[i] = delta[i].detach()\n                        success[i] = True\n                        active[i] = False\n    return success, frozen, step + 1","metadata":{},"outputs":[],"execution_count":null},{"id":"e60d5c38-9bd0-4937-b478-a234a1191b1b","cell_type":"markdown","source":"## 7 · Per-sample geometric bisection, run in lockstep","metadata":{}},{"id":"2d84ab57-5e24-4894-8403-520b420833de","cell_type":"code","source":"def eps50_batch(x, mask, tgt_pad, tgt_len, targets, verbose=True):\n    \"\"\"Per-sample geometric bisection for the smallest budget that still works.\n\n    ---------------------------------------------------------------------------\n    BUG FIXED 2026-09-22 (found in the rung-1 T5 run, 40% censoring).\n\n    The previous version ended with\n\n        ok, best, _ = pgd_batch(..., hi, CONFIRM_STEPS, delta_init=best)\n\n    which REASSIGNED `best` wholesale to pgd_batch's return value. That return\n    is a `frozen` buffer holding zeros for any sample that did not succeed on\n    that particular pass -- so one unlucky final pass wiped out a delta that had\n    already been verified at a larger budget. Those samples were then written\n    out as censored, with SNR 70-86 dB (the signature of delta == 0) even though\n    the ceiling probe had already proved them attackable.\n\n    In the rung-1 run that discarded ~6 of 25 real measurements. Worse, it does\n    so at a RATE THAT VARIES BY MODEL, so comparing median eps50 across rungs\n    would have been comparing differently-biased subsamples -- straight into\n    RQ1's headline.\n\n    Fix: keep a running `best_eps` / `best_delta` that is only ever updated on a\n    VERIFIED success. Nothing overwrites a confirmed result. Censoring now means\n    exactly one thing: the attack failed at every budget tried, ceiling included.\n\n    Side benefit: eps50 becomes a guaranteed UPPER BOUND on the true minimum.\n    Optimiser noise can only make it conservative, never wrong in an\n    unpredictable direction -- which is the property you want in a metric you\n    are about to compare across four models.\n    ---------------------------------------------------------------------------\n    \"\"\"\n    B = x.shape[0]\n    lo = torch.full((B,), EPS_LO, device=DEVICE)\n    hi = torch.full((B,), EPS_HI, device=DEVICE)\n    best_eps   = torch.full((B,), float(\"nan\"), device=DEVICE)   # verified only\n    best_delta = torch.zeros_like(x)\n\n    def record(okt, eps_vec, d):\n        \"\"\"Adopt a result only when it is both successful and better.\"\"\"\n        nonlocal best_eps, best_delta\n        better = okt & (torch.isnan(best_eps) | (eps_vec < best_eps))\n        best_eps   = torch.where(better, eps_vec, best_eps)\n        best_delta = torch.where(better.view(-1, 1), d, best_delta)\n\n    # Ceiling probe. Failure here is genuine right-censoring.\n    ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, hi, CONFIRM_STEPS)\n    okt = torch.tensor(ok, device=DEVICE)\n    record(okt, hi, d)\n    feasible = ok.copy()\n    if verbose:\n        print(f\"  ceiling probe eps={EPS_HI:.1e}: {feasible.sum()}/{B} feasible, \"\n              f\"{(~feasible).sum()} censored\")\n\n    for rnd in range(BISECT_ROUNDS):\n        mid = torch.sqrt(lo * hi)\n        ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, mid,\n                             PROBE_STEPS, delta_init=best_delta)\n        ok = ok & feasible\n        okt = torch.tensor(ok, device=DEVICE)\n        hi = torch.where(okt, mid, hi)\n        lo = torch.where(okt, lo, mid)\n        record(okt, mid, d)\n        if verbose:\n            print(f\"  round {rnd+1}/{BISECT_ROUNDS}: {ok.sum()}/{B} succeeded \"\n                  f\"| median eps {mid.median().item():.3e}\")\n\n    # Budget-consistency round: one more bisection at the FULL step budget, so\n    # the reported boundary is measured at the budget eps50 is defined at rather\n    # than at the cheaper probe budget. See RUNLOG for why this matters.\n    #\n    # Skipped in FIXED_EPS mode: there is no bracket to refine. The ceiling probe\n    # already ran at CONFIRM_STEPS, so the single reported budget is measured at\n    # the full budget and needs no second pass.\n    if BISECT_ROUNDS == 0:\n        eps       = best_eps.cpu().numpy()\n        confirmed = ~np.isnan(eps)\n        if verbose:\n            print(f\"  -> FIXED_EPS {EPS_HI:.1e}: {confirmed.sum()}/{B} succeeded, \"\n                  f\"{(~confirmed).sum()} resisted\")\n        return eps, np.ones(B), confirmed, best_delta\n\n    mid = torch.sqrt(lo * hi)\n    ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, mid,\n                         CONFIRM_STEPS, delta_init=best_delta)\n    ok = ok & feasible\n    okt = torch.tensor(ok, device=DEVICE)\n    record(okt, mid, d)\n    if verbose:\n        print(f\"  budget-consistency round @ {CONFIRM_STEPS} steps: {ok.sum()}/{B}\")\n\n    eps       = best_eps.cpu().numpy()\n    confirmed = ~np.isnan(eps)                 # censored == never succeeded\n    width     = (hi / lo).cpu().numpy()\n    if verbose:\n        print(f\"  -> {confirmed.sum()}/{B} measured, {(~confirmed).sum()} censored\")\n    return eps, width, confirmed, best_delta\n","metadata":{},"outputs":[],"execution_count":null},{"id":"f2dc1406-eb88-4a9d-918b-21eae52c8a6c","cell_type":"markdown","source":"## 8 · Validate the batching before trusting it\n\nBatching is a substantial rewrite. If masked normalisation or the per-sample CTC lengths are\nwrong, results would be quietly biased rather than obviously broken. So: run two utterances\nbatched, then the same two alone, and compare.\n\n> **This costs about three extra eps50 searches (~1–2 GPU-h).** Run it once, for\n> `(RUNG=1, EDIT_TYPE=\"T5\")`, record the agreement in `RUNLOG.md`, then set\n> `VALIDATE_BATCHING = False` for the other seven runs. It is a one-off correctness check on\n> the batching code, and that code does not change between runs.","metadata":{}},{"id":"25253f34-ffc9-4e53-94f0-ad5423628ba4","cell_type":"code","source":"if VALIDATE_BATCHING and len(ev) >= 2:\n    sub = ev.head(2)\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    print(\"batched (2 together):\")\n    t0 = time.time(); e_b, _, c_b, _ = eps50_batch(xb, mb, tpad, tl, tgts); tb = time.time()-t0\n\n    print(\"\\nsingle (one at a time):\")\n    e_s = []\n    t0 = time.time()\n    for i in range(2):\n        xs, ms = xb[i:i+1], mb[i:i+1]\n        e, _, c, _ = eps50_batch(xs, ms, tpad[i:i+1], tl[i:i+1], [tgts[i]])\n        e_s.append(e[0])\n    ts = time.time()-t0\n\n    print(f\"\\n{'':6s} {'batched':>12s} {'single':>12s}  {'rel diff':>9s}\")\n    for i in range(2):\n        a, b = e_b[i], e_s[i]\n        rd = \"n/a\" if (np.isnan(a) or np.isnan(b)) else f\"{100*abs(a-b)/b:8.1f}%\"\n        print(f\"  #{i}  {a:12.3e} {b:12.3e}  {rd:>9s}\")\n    print(f\"\\nwall-clock: batched {tb:.0f}s vs single {ts:.0f}s -> {ts/tb:.2f}x\")\n    print(\"Agreement within ~15% is fine (PGD is stochastic). Larger means a batching bug.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"0a1fa4c8-2792-45d6-890a-310d114fc828","cell_type":"markdown","source":"## 8b · Is eps50 converged at this step budget?\n","metadata":{}},{"id":"70df87bd-2a5b-488d-9f03-718c7b68e39c","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# CONVERGENCE CHECK.\n#\n# eps50 is defined relative to CONFIRM_STEPS. If doubling the budget moves it\n# materially, the metric is still tracking optimiser convergence rather than\n# model robustness -- and since rungs may converge at different rates, that\n# would leak straight into the RQ1 comparison. Three utterances is enough to\n# detect it.\n# ---------------------------------------------------------------------------\nif len(ev) >= 3:\n    sub = ev.head(3)\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    _saved = CONFIRM_STEPS\n    print(f\"at CONFIRM_STEPS = {_saved}:\")\n    e1, _, _, _ = eps50_batch(xb, mb, tpad, tl, tgts, verbose=False)\n    CONFIRM_STEPS = _saved * 2\n    print(f\"at CONFIRM_STEPS = {CONFIRM_STEPS}:\")\n    e2, _, _, _ = eps50_batch(xb, mb, tpad, tl, tgts, verbose=False)\n    CONFIRM_STEPS = _saved\n\n    print(f\"\\n{'utt':>4s} {'@'+str(_saved):>12s} {'@'+str(_saved*2):>12s} {'shift':>9s}\")\n    shifts = []\n    for i in range(len(sub)):\n        a, b = e1[i], e2[i]\n        if np.isnan(a) or np.isnan(b):\n            print(f\"{i:>4d} {'censored':>12s}\"); continue\n        sh = abs(a - b) / a; shifts.append(sh)\n        print(f\"{i:>4d} {a:12.3e} {b:12.3e} {100*sh:8.1f}%\")\n    if shifts:\n        ms = float(np.median(shifts))\n        print(f\"\\nmedian shift on doubling the budget: {100*ms:.1f}%\")\n        if ms < 0.10:\n            print(\"Converged. eps50 is stable at this budget -- safe to compare rungs.\")\n        else:\n            print(\"NOT CONVERGED. Raise CONFIRM_STEPS and re-run, or state explicitly\")\n            print(\"that eps50 is budget-relative and hold the budget fixed across ALL\")\n            print(\"rungs so the comparison stays internally valid.\")\n        PROV[\"convergence_shift_pct\"] = round(100 * ms, 1)\n","metadata":{},"outputs":[],"execution_count":null},{"id":"bcf2ce5e-7059-4471-a637-beb919d074b5","cell_type":"markdown","source":"## 9 · The main sweep","metadata":{}},{"id":"d55f15c4-ed5d-44c7-9027-c0df49f472d2","cell_type":"code","source":"ev = ev.sort_values(\"duration\").reset_index(drop=True)   # length buckets = less padding\nOUT = f\"{WORK}/eps50_rung{RUNG}_{CHECKPOINT}_{EDIT_TYPE}.csv\"\n\n# Resume: if a previous session ran out of time, attach its output dataset as an\n# input and the utterances it already measured are carried over, not redone.\n#\n# Two fixes over NB-05 (2026-09-26):\n#  (a) the old glob was \"/kaggle/input/*/<file>\", only TWO levels deep, while your\n#      own datasets mount at /kaggle/input/datasets/<user>/<name> -- three levels.\n#      So resume silently never fired. Same defect class as the flat-glob bug\n#      already fixed in section 3; it survived here. Use the resolved `inputs`.\n#  (b) the superseded 38-utt runs wrote a CSV with the SAME basename\n#      (eps50_rung1_<ckpt>_T5.csv). Attaching one of those would have pulled 38\n#      foreign utterances into `results` and contaminated the output, because the\n#      preflight gate has already run by this point. Rows are now checked against\n#      the current attack set and rejected if they do not belong to it.\nresults = []\nfor _p in inputs:\n    _f = f\"{_p}/{os.path.basename(OUT)}\"\n    if not os.path.exists(_f):\n        continue\n    prev = pd.read_csv(_f)\n    _foreign = set(prev.utt_id) - set(ev.utt_id)\n    assert not _foreign, (\n        f\"{os.path.basename(_p)} holds {len(prev)} rows, {len(_foreign)} of which are \"\n        f\"NOT in the current 25-utterance attack set. That dataset was almost \"\n        f\"certainly produced on the 38-utt NB-03 set (the superseded \"\n        f\"eps50-rung1-* runs). Detach it -- resuming from it would mix two \"\n        f\"disjoint utterance sets into one CSV.\")\n    results.extend(prev.to_dict(\"records\"))\n    print(f\"resuming: {len(prev)} utterances already measured in {os.path.basename(_p)}\")\ndone = {r[\"utt_id\"] for r in results}\ntodo = ev[~ev.utt_id.isin(done)].reset_index(drop=True)\n\nbatches = [todo.iloc[i:i+BATCH_SIZE] for i in range(0, len(todo), BATCH_SIZE)]\nprint(f\"{len(todo)} utterances remaining -> {len(batches)} batches of up to {BATCH_SIZE}\\n\")\n\nDEADLINE = time.time() + SESSION_HOURS * 3600\nt_start = time.time()\n\nfor bi, sub in enumerate(batches):\n    if time.time() > DEADLINE:\n        print(f\"*** SESSION BUDGET ({SESSION_HOURS} h) reached. \"\n              f\"{sum(len(b) for b in batches[bi:])} utterances not yet measured. ***\\n\")\n        break\n\n    print(f\"=== batch {bi+1}/{len(batches)} ({len(sub)} utts, \"\n          f\"{sub.duration.min():.1f}-{sub.duration.max():.1f}s) ===\")\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    t0 = time.time()\n    eps, width, conf, best = eps50_batch(xb, mb, tpad, tl, tgts)\n    dt = time.time() - t0\n\n    x_adv = int16_rt(torch.clamp(xb + best, -1, 1)) * mb\n    snrs = snr_db_b(xb, x_adv, mb)\n    hyp = decode_b(x_adv, mb)\n\n    for i, (_, r) in enumerate(sub.iterrows()):\n        results.append(dict(\n            utt_id=r[\"utt_id\"], rung=RUNG, edit_type=EDIT_TYPE, checkpoint=CHECKPOINT,\n            duration=r[\"duration\"], source_text=r[\"text\"], target_text=r[\"target_text\"],\n            source_tokens=int(r[\"source_tokens\"]), target_tokens=int(r[\"target_tokens\"]),\n            eps50=float(eps[i]) if not np.isnan(eps[i]) else None,\n            # SNR/decode are meaningful only where an attack was verified.\n            # Writing them for censored rows produced the 70-86 dB values in\n            # the rung-1 run, which looked like data but were delta == 0.\n            censored=bool(np.isnan(eps[i])),\n            eps50_per_token=(float(eps[i]) / r[\"target_tokens\"]) if not np.isnan(eps[i]) else None,\n            bracket_width=float(width[i]), snr_db=float(snrs[i]),\n            decoded=(hyp[i] if not np.isnan(eps[i]) else None),\n            exact_match=(bool(hyp[i] == r[\"target_text\"])\n                         if not np.isnan(eps[i]) else False),\n            use_amp=USE_AMP, model_dir=MODEL_DIR,\n        ))\n\n    # Write after EVERY batch. The original only wrote at the very end, so a\n    # session killed on the last batch threw away hours of finished work.\n    pd.DataFrame(results).to_csv(OUT, index=False)\n\n    remaining = len(batches) - (bi + 1)\n    print(f\"  batch done in {dt/60:.1f} min | {(~np.isnan(eps)).sum()}/{len(sub)} uncensored \"\n          f\"| {len(results)} rows written\")\n    if remaining:\n        print(f\"  projected for the {remaining} remaining batches: \"\n              f\"{remaining*dt/3600:.1f} h \"\n              f\"({'fits' if time.time()+remaining*dt < DEADLINE else 'WILL NOT FIT this session'})\\n\")\n    else:\n        print()\n\nprint(f\"TOTAL {(time.time()-t_start)/60:.1f} min this session | {len(results)}/{len(ev)} measured\")","metadata":{},"outputs":[],"execution_count":null},{"id":"ca4275c8-c04e-4a05-8a0d-2a022bfca4b5","cell_type":"markdown","source":"## 10 · Results and save","metadata":{}},{"id":"3d039dbc-c7f9-4d0b-ae7a-103beee3f748","cell_type":"code","source":"CK_TAG = \"final\" if CHECKPOINT == \"final\" else \"ms\" + CHECKPOINT.split(\"-\")[1]\nif FIXED_EPS is not None:\n    CK_TAG += \"-fixedeps\"        # keep trajectory runs separate from eps50 runs\n# Every checkpoint of a rung gets its OWN dataset. Before this, 100h final and\n# all three 100h milestones were told to publish as \"eps50-rung100-T5\" (the old\n# NB-05 naming), so \"New Version\" on that dataset silently replaced one\n# checkpoint's results with another's.\n#\n# NB-05b uses an \"eps50b-\" prefix so these 25-utterance runs land in NEW datasets\n# and cannot version over the four superseded 38-utterance \"eps50-rung1-*\" ones,\n# which are kept for the record. Choose \"New Dataset\" when publishing, not\n# \"New Version\".\ndf = pd.DataFrame(results)\nok = df[~df.censored] if len(df) else df\ncomplete = len(df) >= len(ev)\n\n# Clean WER on the SAME attack utterances, measured here so a trajectory run\n# needs no NB-02 pass. The figure plots attack success against clean WER and both\n# numbers must come from the same checkpoint.\nif FIXED_EPS is not None:\n    try:\n        import jiwer\n        _wav, _msk = load_batch(ev)\n        _hyp = decode_b(_wav, _msk)\n        _ref = [norm_bn(t) for t in ev.text]\n        _pr = [(r, h) for r, h in zip(_ref, _hyp) if r.strip()]\n        PROV[\"clean_wer_attackset\"] = float(jiwer.wer([p[0] for p in _pr], [p[1] for p in _pr]))\n        PROV[\"clean_cer_attackset\"] = float(jiwer.cer([p[0] for p in _pr], [p[1] for p in _pr]))\n        print(f\"clean WER on the attack set: {PROV['clean_wer_attackset']:.4f}  \"\n              f\"CER {PROV['clean_cer_attackset']:.4f}   <- x-axis for the trajectory figure\")\n        del _wav, _msk; torch.cuda.empty_cache()\n    except Exception as e:\n        print(f\"clean-WER measurement failed ({type(e).__name__}: {e}); \"\n              f\"fall back to NB-02's clean_wer_all_checkpoints.csv\")\n\nprint(f\"=== RUNG {RUNG}h | {EDIT_TYPE}\"\n      + (f\" | FIXED_EPS {FIXED_EPS:.1e} ===\" if FIXED_EPS is not None else \" ===\"))\nif FIXED_EPS is not None:\n    print(\"  FIXED_EPS mode: every success is recorded at exactly the fixed budget, so\")\n    print(\"  eps50, its IQR and its bootstrap CI below are DEGENERATE and carry no\")\n    print(\"  information. Use ASR@eps and the per-utterance `censored` column only.\")\nprint(f\"utterances      : {len(df)} of {len(ev)}\")\nprint(f\"censored        : {df.censored.sum()} ({100*df.censored.mean():.0f}%)  <- report this\")\nif len(ok):\n    print(f\"median eps50    : {ok.eps50.median():.3e}  ({ok.eps50.median()*32767:.1f} int16 units)\")\n    print(f\"IQR             : {ok.eps50.quantile(.25):.3e} - {ok.eps50.quantile(.75):.3e}\")\n    print(f\"per target token: {ok.eps50_per_token.median():.3e}   <- length-normalised\")\n    print(f\"median SNR      : {ok.snr_db.median():.1f} dB\")\n    print(f\"exact match     : {ok.exact_match.sum()}/{len(ok)}\")\n\n    # A bracket width near 1.0 with eps50 at the floor means the search hit\n    # EPS_LO: LEFT-censored, and just as biasing as the right-censored kind.\n    at_floor = int(((ok.eps50 <= EPS_LO * 1.05)).sum())\n    if at_floor:\n        print(f\"WARNING: {at_floor} utterances sit at the EPS_LO floor ({EPS_LO:.0e}) \"\n              \"-- left-censored. Lower EPS_LO and re-run this rung.\")\n\n    b = [ok.eps50.sample(len(ok), replace=True).median() for _ in range(2000)]\n    print(f\"bootstrap 95% CI: [{np.percentile(b,2.5):.3e}, {np.percentile(b,97.5):.3e}]\")\n\ndf.to_csv(f\"{WORK}/eps50_rung{RUNG}_{CHECKPOINT}_{EDIT_TYPE}.csv\", index=False)\nPROV.update(n_utts=len(df), n_expected=len(ev), complete=bool(complete),\n            n_censored=int(df.censored.sum()) if len(df) else 0,\n            median_eps50=float(ok.eps50.median()) if len(ok) else None,\n            median_eps50_per_token=float(ok.eps50_per_token.median()) if len(ok) else None,\n            median_snr=float(ok.snr_db.median()) if len(ok) else None,\n            minutes=round((time.time()-t_start)/60, 1))\njson.dump(PROV, open(f\"{WORK}/_provenance.json\", \"w\"), indent=2)\n\nKAGGLE_USER = os.environ.get(\"KAGGLE_USERNAME\") or \"YOURUSERNAME\"\njson.dump({\"title\": f\"eps50b-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}\",\n           \"id\": f\"{KAGGLE_USER}/eps50b-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}\",\n           \"licenses\": [{\"name\": \"CC0-1.0\"}]},\n          open(f\"{WORK}/dataset-metadata.json\", \"w\"), indent=2)\n\nif not complete:\n    print(\"\\n\" + \"=\" * 70)\n    print(f\"  PARTIAL: {len(df)}/{len(ev)} utterances measured. Nothing is lost.\")\n    print(f\"    1. Output panel -> Create Dataset (or New Version) \"\n          f\"-> eps50b-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}\")\n    print( \"    2. Attach that dataset back as an INPUT to this notebook\")\n    print( \"    3. Re-commit. It skips what is already done.\")\n    print(\"=\" * 70)\nelse:\n    print(f\"\\nRUNG {RUNG}h / {EDIT_TYPE} COMPLETE -- publish as \"\n          f\"eps50b-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()} and record the median in RUNLOG.md.\")\n\nprint(\"\\n\" + json.dumps(PROV, indent=2))\ndf.head(10)\n\n# ---------------------------------------------------------------------------\n# CO-PRIMARY METRIC: attack success rate at a FIXED budget.\n#\n# Median eps50 is computed only over UNCENSORED utterances, so if the censoring\n# rate differs between rungs -- and it will -- the medians describe different\n# subpopulations and are not directly comparable. That is a censoring bias\n# sitting directly under RQ1's headline.\n#\n# ASR@ceiling has no such problem: every utterance contributes, success or\n# failure, so it is comparable across rungs by construction.\n#\n# The two measure different things and should both be reported:\n#   ASR@ceiling -> can this model be STEERED to an exact target at all?\n#   median eps50 -> when it can, how much perturbation does it take?\n#\n# Expect them to move in opposite directions. A better model is usually easier\n# to aim (lower censoring) but harder to push (higher eps50). Reporting only\n# one of them would tell half the story, and the flattering half depends on\n# which one you pick.\n# ---------------------------------------------------------------------------\nASR_AT_CEILING = float((~df.censored).mean())\nprint(f\"\\nASR@eps={EPS_HI:.0e}   : {ASR_AT_CEILING:.1%}  ({(~df.censored).sum()}/{len(df)})\")\nprint(f\"  <- censoring-free, uses all {len(df)} utterances, comparable across rungs\")\nprint(f\"median eps50 is over the {(~df.censored).sum()} steerable utterances only\")\n\n# ---------------------------------------------------------------------------\n# LEFT-censoring: the mirror image of the block above, and the one this notebook\n# had no check for. If eps50 lands at the bottom of the bracket, the attack\n# succeeded at the smallest budget searched, so the value is an UPPER BOUND and\n# the true eps50 is somewhere below it. Reporting it as a point estimate\n# understates how cheap the attack is -- in the direction that flatters RQ2.\n# ---------------------------------------------------------------------------\nfloor_hits = int((ok.eps50 <= EPS_LO * 1.15).sum()) if len(ok) else 0\nPROV[\"left_censored\"] = floor_hits\nif floor_hits:\n    print(f\"\\n*** LEFT-CENSORED: {floor_hits}/{len(ok)} eps50 values sit at the \"\n          f\"bracket floor ({EPS_LO:.1e}) ***\")\n    print(\"    These are upper bounds, not measurements. Lower EPS_LO by a decade\")\n    print(\"    and re-run, or report them as censored-below and say so in the paper.\")\n\nif len(ok):\n    b = [ok.eps50.sample(len(ok), replace=True).median() for _ in range(2000)]\n    lo_ci, hi_ci = np.percentile(b, 2.5), np.percentile(b, 97.5)\n    print(f\"\\neps50 bootstrap 95% CI spans {hi_ci/lo_ci:.2f}x\")\n    if hi_ci / lo_ci > 1.8:\n        print(\"  WIDE. Between-rung differences must exceed this to be credible.\")\n        print(\"  If the ladder effect is smaller than the CI, raise CONFIRM_STEPS or\")\n        print(\"  run each condition twice -- but only if the effect turns out marginal.\")\n\nPROV[\"asr_at_ceiling\"] = round(ASR_AT_CEILING, 4)\nPROV[\"eps50_ci_ratio\"] = round(float(hi_ci / lo_ci), 3) if len(ok) else None\n","metadata":{},"outputs":[],"execution_count":null},{"id":"8ee9f775-1883-482b-b592-002224ba63aa","cell_type":"markdown","source":"---\n\n## Record this in RUNLOG.md, then run the next combination\n\n| Rung | T1 median eps50 | T1 censored | T5 median eps50 | T5 censored |\n|---|---|---|---|---|\n| 1 h | | | | |\n| 5 h | | | | |\n| 25 h | | | | |\n| 100 h | | | | |\n\n### What you are looking for\n\n**RQ1:** eps50 should *rise* with training hours — more data, more perturbation needed. That is\nthe headline. It only counts after NB-02's iso-WER adjustment, since clean WER is a covariate.\n\n**RQ2:** T1 should cost less than T5 at matched target length. Compare the\n`eps50_per_token` column, not the raw one.\n\n### Watch the censoring rate\n\nIf the 1 h rung shows **high censoring**, that is the paradox the plan predicted: a model whose\noutput is already unstable can be easy to *disrupt* yet hard to steer to a *specific* target.\nThat is a finding, not a failure. Report censoring alongside eps50 and use Kaplan–Meier rather\nthan silently dropping those rows.\n\n### Next\n**NB-06 — the LM-weight sweep.** Cheapest experiment in the thesis and the one that turns the\ncorrelation into a mechanism. Prioritise it over everything except the ladder itself.","metadata":{}}]}