{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13"},"accelerator":"GPU"},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"94a58e8e-4db8-478a-993c-2d3a93500cf2","cell_type":"markdown","source":"# NB-05 — eps50 Across the Data Ladder (batched)\n\n**This notebook produces the headline result of the thesis.**\n\nRun it once per `(RUNG, EDIT_TYPE)` pair: 4 rungs × {T1, T5} = 8 committed versions.\n\n---\n\n### Why this is a rewrite, not a copy of NB-04\n\nNB-04 measured **17.6 min/utterance**. Scaled up that was ~94 GPU-hours against a 72-hour\ntotal budget. Five changes bring it to roughly **8 hours** for the whole sweep:\n\n| Change | Gain | Why |\n|---|---|---|\n| **Batch 8 utterances** | **4–5×** | A T4 attacking one 7 s clip is mostly idle. Biggest win by far. |\n| Bracket `[1e-4, 2e-3]` | ~1.4× | NB-04 found eps50 = 3.7e-4. The old `[1e-5, 2e-2]` wasted iterations in dead regions. |\n| 300-step probes + confirm | ~2.5× | Failures cost the full step budget; most are decidable early. |\n| Stall abort | ~1.5× | No CTC improvement in 100 steps means it will not converge. |\n| 5 bisection rounds, not 7 | 1.4× | Beyond 5, precision is spurious. Report bracket width as uncertainty. |\n\n**AMP is enabled.** NB-04b's seed-variance control settled it: AMP reproduced fp32's eps50\n*exactly* (0.0% apart) while fp32 disagreed with *itself* by 8.9% across seeds, with zero\nsteady-state bad gradients and a 2.4x search speedup. Combined with batching, the full\n8-condition sweep should land near **4-5 GPU-hours** rather than the original 164.\n\n### Kaggle settings\n| Setting | Value |\n|---|---|\n| Accelerator | **GPU T4 ×2** |\n| Internet | On |\n| Persistence | Variables and Files |\n| **Inputs** | 1. `bn-asr-splits` 2. `bengaliai-speech` 3. `ckpt-xlsr-{RUNG}h` |\n| Output | `eps50-rung{RUNG}-{EDIT_TYPE}` |","metadata":{}},{"id":"ffc72064-f3b3-4df5-ab76-80b956bfd280","cell_type":"markdown","source":"## 1 · Configuration","metadata":{}},{"id":"804ecdd2-1388-4836-9eb2-85a4e2a877c5","cell_type":"code","source":"RUNG        = 25          # <<<< 1, 5, 25, 100   (hours of AUDIO the model was trained on)\nEDIT_TYPE   = \"T5\"       # <<<< \"T1\" (negation flip) or \"T5\" (full retarget)\n\n# Which checkpoint of that rung. The iso-WER pairs from NB-02 need MILESTONES,\n# not just finals: a 100h model caught mid-training, at the moment its accuracy\n# matched a low-data model's. Attacking only the four finals cannot test RQ1's\n# strong form.\n#   \"final\" | \"milestone-1000\" | \"milestone-2000\" | \"milestone-3500\"\nCHECKPOINT  = \"milestone-3500\"\n\nN_EVAL      = 25         # frozen attack set size (cut from 40 on budget grounds)\n\n# The NB-03 corpus is REQUIRED, not optional. Two reasons, both learned the hard\n# way on 2026-09-25: (a) attack_eval_25.csv contains 0 utterances with a standalone\n# \"না\", so T1 cannot run on it at all; (b) a paired T1-vs-T5 comparison needs BOTH\n# edit types on the SAME audio, and silently falling back to auto-generated targets\n# produced a T5 run on the wrong 25 utterances that looked successful.\nREQUIRE_CORPUS = False\nBATCH_SIZE  = 8          # per-batch utterances; drop to 4 if you hit OOM\nUSE_AMP     = True       # <<<< NB-04b run 2 (2026-09-20), which added the fp32\n                         #      seed-variance CONTROL the first run lacked:\n                         #        fp32 seed A   eps50 5.393e-04\n                         #        fp32 seed B   eps50 4.911e-04  -> 8.9% noise floor\n                         #        AMP           eps50 5.393e-04  -> 0.0% from seed A\n                         #        0 steady-state bad grads, 2.40x faster\n                         #      Run 1 rejected AMP on \"6 bad gradients > 5\". That\n                         #      threshold was arbitrary and those 6 were GradScaler\n                         #      calibrating (it halves from 65536 until no overflow).\n                         #      Superseded -- do not revert to False.\n\n# Hard wall-clock stop, same reason as NB-01: Kaggle kills a committed GPU\n# session at 12 h. Results are written to CSV after EVERY batch, so a session\n# that runs out of time loses nothing -- publish the partial CSV, attach it\n# back as an input, re-commit, and it skips the utterances already measured.\nSESSION_HOURS = 8.0\n\n# T1 targets are the source sentence minus one word, so they are FAR cheaper\n# than T5 and land at the bottom of the T5 bracket. The 2026-09-25 T1 probe on\n# 100h-final returned 1.048e-04 against an EPS_LO of 1.000e-04: every value\n# pinned to the floor, i.e. an UPPER BOUND, not a measurement. T1 needs one more\n# decade of headroom, and one more bisection round to keep the resolution.\n#   T5: [1e-4, 2e-3], 5 rounds -> ~10% resolution (unchanged, so T5 stays comparable)\n#   T1: [1e-5, 2e-3], 6 rounds -> ~9%  resolution\n# FIXED_EPS mode. Set to a float to skip bisection entirely and measure only\n# \"does the attack succeed at this budget\". Costs roughly a sixth of a full\n# eps50 search, because it runs one PGD at one epsilon instead of 1 + rounds + 1.\n#\n# Purpose: the trajectory figure that breaks the early-checkpoint confound.\n# In both significant iso-WER pairs the robust model is an EARLY CHECKPOINT of a\n# long run and the fragile one a CONVERGED model of a short run, so \"robustness\n# from earliness\" is not excluded by the compute control. Attacking several\n# checkpoints along EACH rung at one fixed epsilon and plotting attack success\n# against clean WER makes earliness vary WITHIN every trajectory, so it cannot\n# explain separation BETWEEN trajectories.\n#\n# Milestones exist for rungs 25 and 100 only (NB-01: MILESTONE_STEPS is empty\n# below rung 25), so the two available trajectories are those.\nFIXED_EPS = 2e-3         # <<<< e.g. 2e-3 for the trajectory runs; None = full eps50\n\nEPS_LO, EPS_HI = (1e-5, 2e-3) if EDIT_TYPE == \"T1\" else (1e-4, 2e-3)\nBISECT_ROUNDS  = 6 if EDIT_TYPE == \"T1\" else 5\nif FIXED_EPS is not None:\n    EPS_HI = float(FIXED_EPS)\n    EPS_LO = min(EPS_LO, EPS_HI)\n    BISECT_ROUNDS = 0\nPROBE_STEPS    = 400     # steps per bisection probe (matches NB-04b)\nCONFIRM_STEPS  = 800     # steps for the final confirmation at the chosen eps\nSTALL_PATIENCE = 100     # abort a probe after this many steps with no loss improvement\nCHECK_EVERY    = 50\nLR             = 1e-4\n\n# Costs roughly THREE extra eps50 searches (~1-2 GPU-h) because it runs the\n# same two utterances batched and then one at a time. Run it ONCE, for\n# (RUNG=1, EDIT_TYPE=\"T5\"), record the agreement in RUNLOG.md, then set it\n# False for the remaining seven runs. It is a one-off correctness check on the\n# batching, not something each run needs to repeat.\nVALIDATE_BATCHING = False   # <<<< set True for ONE run only, then back to False.\n                            # NOTE: RUNLOG has no row recording the batching\n                            # agreement figure. It was almost certainly measured\n                            # on the first run (RUNG=1, T5 -- the old defaults),\n                            # but check that run's log and add the number to\n                            # RUNLOG before the thesis cites batched eps50.","metadata":{},"outputs":[],"execution_count":null},{"id":"77ff4194-2f11-46a1-994c-0a0450966fa8","cell_type":"markdown","source":"## 2 · Boilerplate","metadata":{}},{"id":"6c6663ba-4cf5-4ead-8e03-ed54698b24bb","cell_type":"code","source":"import os\nos.environ[\"CUDA_VISIBLE_DEVICES\"] = \"0\"   # attacks optimise the INPUT: never DataParallel\nimport sys, json, time, random, re, glob, unicodedata\nimport numpy as np, pandas as pd, torch\nimport torch.nn.functional as F\ntorch.backends.cudnn.benchmark = True\n\nSEED = 1337\nrandom.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)\nWORK, TEMP, DEVICE = \"/kaggle/working\", \"/kaggle/temp\", \"cuda\"\nassert torch.cuda.is_available(), \"Settings -> Accelerator -> GPU T4 x2\"\n\nPROV = {\"notebook\": \"NB-05-eps50-batched\", \"rung\": RUNG, \"edit_type\": EDIT_TYPE,\n        \"n_eval\": N_EVAL, \"use_amp\": USE_AMP, \"seed\": SEED,\n        \"bracket\": [EPS_LO, EPS_HI], \"bisect_rounds\": BISECT_ROUNDS,\n        \"gpu\": torch.cuda.get_device_name(0), \"torch\": torch.__version__,\n        \"timestamp\": time.strftime(\"%Y-%m-%d %H:%M:%S\")}\nprint(json.dumps(PROV, indent=2))","metadata":{},"outputs":[],"execution_count":null},{"id":"f0d5c04c-b1ce-4e9c-adf8-f8934203908d","cell_type":"code","source":"# See NB-01 for the full explanation: pinning an old `accelerate`/`datasets`\n# silently downgrades numpy mid-session and breaks pandas much later with\n# \"No module named 'numpy.rec'\". Pin numpy to what is already loaded.\n# `datasets` is not imported by this notebook, so it is not installed.\nimport numpy, subprocess, sys\nNP_BEFORE = numpy.__version__\n\n!pip install -q transformers==4.44.2 librosa soundfile jiwer \"numpy=={NP_BEFORE}\" 2>&1 | tail -2\n\nNP_DISK = subprocess.run([sys.executable, \"-c\", \"import numpy; print(numpy.__version__)\"],\n                         capture_output=True, text=True).stdout.strip()\nassert NP_BEFORE == NP_DISK, f\"numpy moved {NP_BEFORE} -> {NP_DISK}; relax the pin that forced it.\"\nprint(f\"numpy {NP_BEFORE} stable\")\n\nfrom transformers import Wav2Vec2ForCTC, Wav2Vec2Processor\nimport transformers, librosa\nprint(\"transformers\", transformers.__version__)\nPROV[\"transformers\"] = transformers.__version__\nPROV[\"numpy\"] = NP_BEFORE\n","metadata":{},"outputs":[],"execution_count":null},{"id":"90f4262a-409e-4fa2-8687-b49a99859742","cell_type":"markdown","source":"## 3 · Load the ladder checkpoint and the frozen attack set","metadata":{}},{"id":"19a045c3-639a-418a-8a57-f0629d5f29c6","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Kaggle does not use one mount layout. Depending on the kernel you get\n#     /kaggle/input/<name>                (older kernels)\n#     /kaggle/input/datasets/<name>       (your own datasets, current)\n#     /kaggle/input/competitions/<name>   (competition data)\n# A flat glob only sees the wrapper dirs -- NB-01 failed with exactly this,\n# reporting: Visible inputs: ['competitions', 'datasets']\n# ---------------------------------------------------------------------------\ninputs = sorted({p for pat in (\"/kaggle/input/*\", \"/kaggle/input/*/*\",\n                               \"/kaggle/input/*/*/*\")\n                 for p in glob.glob(pat) if os.path.isdir(p)})\n\n# Match the rung EXACTLY. The old test was `f\"{RUNG}h\" in p`, and \"5h\" is a\n# substring of \"ckpt-xlsr-25h\" -- with both datasets attached, RUNG=5 could\n# silently attack the 25 h model and corrupt the headline figure.\ncands = [p for p in inputs if re.search(rf\"ckpt-xlsr-{RUNG}h$\", os.path.basename(p))]\nassert len(cands) == 1, (\n    f\"Need exactly one input matching 'ckpt-xlsr-{RUNG}h', found {len(cands)}: \"\n    f\"{[os.path.basename(c) for c in cands]}\\n\"\n    f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\nCKPT = cands[0]\nif CHECKPOINT == \"final\":\n    MODEL_DIR = CKPT if os.path.exists(f\"{CKPT}/config.json\") else f\"{CKPT}/final\"\nelse:\n    MODEL_DIR = f\"{CKPT}/milestones/step-{CHECKPOINT.split('-')[1]}\"\n    assert os.path.isdir(MODEL_DIR), (\n        f\"{MODEL_DIR} missing. Milestones exist only for rungs >= 25h \"\n        f\"(NB-01 saved them at steps 1000/2000/3500).\")\nassert os.path.exists(f\"{MODEL_DIR}/config.json\"), \\\n    f\"no config.json under {MODEL_DIR} -- did NB-01 finish and publish this rung?\"\nprint(\"checkpoint:\", MODEL_DIR)\n\nprocessor = Wav2Vec2Processor.from_pretrained(MODEL_DIR)\nmodel = Wav2Vec2ForCTC.from_pretrained(MODEL_DIR).float().to(DEVICE).eval()\nfor p in model.parameters(): p.requires_grad_(False)\nmodel.config.apply_spec_augment = False      # no masking noise during attacks\n\nDO_NORMALIZE = bool(getattr(processor.feature_extractor, \"do_normalize\", True))\nBLANK_ID = model.config.pad_token_id\nprint(f\"do_normalize {DO_NORMALIZE} | blank {BLANK_ID} | vocab {model.config.vocab_size}\")\n\nSPLITS = next((p for p in inputs if \"splits\" in p.lower()), None)\nassert SPLITS, (f\"Attach bn-asr-splits. Visible: {[os.path.basename(p) for p in inputs]}\")\n# Which utterances to attack. The NB-03 set is used for BOTH edit types when it\n# is attached: T1 vs T5 is only a paired comparison if both run on the same audio.\n# attack_eval_25.csv has 0 standalone-\"না\" utterances (verified 2026-09-25), so T1\n# can never use it.\nEV_CSV = f\"{SPLITS}/attack_eval_{N_EVAL}.csv\"\n_alt = next((f\"{p}/attack_eval_t1.csv\" for p in inputs\n             if os.path.exists(f\"{p}/attack_eval_t1.csv\")), None)\nif _alt:\n    EV_CSV = _alt\n    print(f\"NB-03 attack set: {EV_CSV}\")\nelse:\n    assert EDIT_TYPE != \"T1\" and not REQUIRE_CORPUS, (\n        f\"attack_eval_t1.csv not found and EDIT_TYPE={EDIT_TYPE!r}. Attach \"\n        f\"bn-attack-targets (run NB-03 first).\\n\"\n        f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n    print(f\"falling back to the frozen T5 set: {EV_CSV}\")\nev = pd.read_csv(EV_CSV)\nPROV[\"attack_eval_csv\"] = os.path.basename(EV_CSV)\nPROV[\"fixed_eps\"] = FIXED_EPS\nPROV[\"bisect_rounds\"] = BISECT_ROUNDS\n\n# ---------------------------------------------------------------------------\n# Re-resolve the audio root, same hazard as NB-01: NB-00 baked in the mount\n# point from ITS session. Here the whole attack set is only 25 files, so check\n# every one of them rather than sampling.\n# ---------------------------------------------------------------------------\ndef resolve_audio(df):\n    p0 = str(df[\"path\"].iloc[0])\n    if os.path.exists(p0):\n        return df, os.path.dirname(p0)\n    stem, ext = os.path.basename(os.path.dirname(p0)), os.path.splitext(p0)[1]\n    cands = [d for pat in (f\"/kaggle/input/*/{stem}\",\n                           f\"/kaggle/input/*/*/{stem}\",\n                           f\"/kaggle/input/*/*/*/{stem}\")\n             for d in glob.glob(pat) if os.path.isdir(d)]\n    assert cands, (f\"Audio directory '{stem}' not found under /kaggle/input.\\n\"\n                   \"Attach: + Add Input -> Competitions -> 'bengaliai speech'.\\n\"\n                   f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n    df = df.copy()\n    df[\"path\"] = cands[0] + \"/\" + df[\"utt_id\"].astype(str) + ext\n    return df, cands[0]\n\nev, AUDIO_ROOT = resolve_audio(ev)\nabsent = [p for p in ev[\"path\"] if not os.path.exists(p)]\nassert not absent, f\"{len(absent)}/{len(ev)} attack audio files missing, e.g.\\n  {absent[0]}\"\nPROV[\"model_dir\"] = MODEL_DIR; PROV[\"audio_root\"] = AUDIO_ROOT\nprint(f\"audio root: {AUDIO_ROOT}\")\nprint(f\"attack set: {len(ev)} utterances, all present\")","metadata":{},"outputs":[],"execution_count":null},{"id":"07bf8011-837a-48d3-91e4-8b4f31ff6cb6","cell_type":"markdown","source":"## 4 · Build the targets\n\nIf NB-03's hand-built corpus is attached, use it. Otherwise generate automatically:\n\n* **T1** — delete `না` from the reference transcript. Post-verbal, short, unstressed; one\n  syllable inverts the meaning.\n* **T5** — a *different real Bengali sentence* drawn from the corpus with a **matched token\n  count**. This matters: NB-04's T5 target was 9 tokens against a ~55-character source, and\n  forcing a much shorter output is plausibly easier, which would flatter T5 and bias against\n  the T1 hypothesis. Matching length removes that.","metadata":{}},{"id":"74c2dbf7-bbbf-4b0c-a0ec-c51a18f4d32c","cell_type":"code","source":"ZW = dict.fromkeys(map(ord, \"‌‍\"), None)\ndef norm_bn(s):\n    return re.sub(r\"\\s+\", \" \", unicodedata.normalize(\"NFC\", str(s)).translate(ZW)).strip()\ndef ntok(s): return len(processor.tokenizer(norm_bn(s)).input_ids)\n\n# Use the resolved `inputs` list from the cell above, NOT a flat glob. Your own\n# datasets mount at /kaggle/input/datasets/<user>/<name>, three levels deep, so\n# glob(\"/kaggle/input/*\") sees only the wrapper dirs. That is the same failure\n# NB-01 hit, and on 2026-09-25 it silently hid an attached bn-attack-targets:\n# T5 then auto-generated targets on the wrong utterance set and still \"succeeded\".\nTGT_CSV = next((f\"{p}/bn_attack_targets.csv\" for p in inputs\n                if os.path.exists(f\"{p}/bn_attack_targets.csv\")), None)\nassert TGT_CSV or not REQUIRE_CORPUS, (\n    \"bn_attack_targets.csv not found. Attach bn-attack-targets (run NB-03 first), \"\n    f\"or set REQUIRE_CORPUS=False to auto-generate.\\n\"\n    f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n\nif TGT_CSV:\n    tg = pd.read_csv(TGT_CSV)\n    tg = tg[tg.edit_type == EDIT_TYPE]\n    ev = ev.merge(tg[[\"source_utterance_id\", \"target_text\"]],\n                  left_on=\"utt_id\", right_on=\"source_utterance_id\", how=\"inner\")\n    print(f\"using NB-03 hand-built corpus: {len(ev)} pairs\")\nelse:\n    print(\"NB-03 corpus not attached -> generating targets automatically.\")\n    print(\"REPLACE THIS with the hand-built corpus before the final run.\\n\")\n    pool = pd.read_csv(f\"{SPLITS}/test.csv\")\n    pool[\"ntok\"] = pool[\"text\"].map(ntok)\n    ev[\"src_ntok\"] = ev[\"text\"].map(ntok)\n\n    rows = []\n    rng = np.random.RandomState(SEED)\n    for _, r in ev.iterrows():\n        srct = norm_bn(r[\"text\"])\n        if EDIT_TYPE == \"T1\":\n            # STANDALONE particle only. A substring test is wrong: \"না\" occurs\n            # inside many common Bengali words -- নাম (name), নারী (woman),\n            # জানা (to know) -- and str.replace() would carve syllables out of\n            # them, producing a target that is not a negation inversion at all.\n            # Remove the LAST standalone \"না\", which is the post-verbal negation.\n            toks = srct.split()\n            if \"না\" not in toks: continue\n            j = len(toks) - 1 - toks[::-1].index(\"না\")\n            tgt = norm_bn(\" \".join(toks[:j] + toks[j+1:]))\n            if tgt == srct or not tgt: continue\n        else:\n            cand = pool[(pool.utt_id != r[\"utt_id\"]) &\n                        (pool.ntok.between(r[\"src_ntok\"] - 2, r[\"src_ntok\"] + 2))]\n            if len(cand) == 0: continue\n            tgt = norm_bn(cand.sample(1, random_state=int(rng.randint(1e6))).iloc[0][\"text\"])\n        rows.append({**r.to_dict(), \"target_text\": tgt})\n    ev = pd.DataFrame(rows)\n    if EDIT_TYPE == \"T1\":\n        # NB-00 froze the attack set using a SUBSTRING test for \"না\", so the\n        # advertised \"15 with না\" overcounts utterances that actually carry the\n        # standalone negation particle. Find out before spending GPU.\n        print(f\"\\nT1: {len(ev)} of {N_EVAL} utterances carry a standalone না\")\n        assert len(ev) >= 10, (\n            f\"only {len(ev)} T1-attackable utterances -- too few for RQ2. \"\n            f\"Either build the NB-03 corpus, or freeze an additional T1-specific \"\n            f\"attack subset from the test split (without touching bn-asr-splits).\")\n\nev[\"target_text\"] = ev[\"target_text\"].map(norm_bn)\nev[\"target_tokens\"] = ev[\"target_text\"].map(ntok)\nev[\"source_tokens\"] = ev[\"text\"].map(ntok)\nev = ev.reset_index(drop=True)\nPROV[\"n_attackable\"] = int(len(ev))\nPROV[\"n_eval_frozen\"] = int(N_EVAL)\nPROV[\"mean_target_tokens\"] = float(ev.target_tokens.mean())\nPROV[\"mean_source_tokens\"] = float(ev.source_tokens.mean())\nprint(f\"{EDIT_TYPE}: {len(ev)} attackable utterances\")\nprint(f\"target tokens  mean {ev.target_tokens.mean():.1f}  vs source {ev.source_tokens.mean():.1f}\")\nprint(\"  (these should be close for T5; T1 is by construction slightly shorter)\")\nev[[\"utt_id\", \"source_tokens\", \"target_tokens\", \"text\", \"target_text\"]].head()","metadata":{},"outputs":[],"execution_count":null},{"id":"8af8c2e7-b062-4a8a-b000-64a0134b5ea5","cell_type":"markdown","source":"## 5 · Batched differentiable front-end and loss","metadata":{}},{"id":"7f46d9cb-78c1-4f72-a24b-9f5fe5552367","cell_type":"code","source":"SR = 16000\n\ndef load_batch(rows):\n    \"\"\"Length-bucketed batch. Returns padded waveforms + a validity mask.\"\"\"\n    wavs = []\n    for _, r in rows.iterrows():\n        a, _ = librosa.load(r[\"path\"], sr=SR, mono=True)\n        a = np.asarray(a, np.float32)\n        pk = np.abs(a).max()\n        wavs.append(a / max(pk, 1.0))\n    T = max(len(w) for w in wavs)\n    x = torch.zeros(len(wavs), T, dtype=torch.float32)\n    m = torch.zeros(len(wavs), T, dtype=torch.float32)\n    for i, w in enumerate(wavs):\n        x[i, :len(w)] = torch.from_numpy(w); m[i, :len(w)] = 1.0\n    return x.to(DEVICE), m.to(DEVICE)\n\n\ndef frontend_b(x, mask):\n    \"\"\"Masked zero-mean/unit-variance. Computing stats over the PADDING would\n    corrupt short utterances in a mixed-length batch.\"\"\"\n    if DO_NORMALIZE:\n        n = mask.sum(-1, keepdim=True).clamp(min=1)\n        mu = (x * mask).sum(-1, keepdim=True) / n\n        var = (((x - mu) * mask) ** 2).sum(-1, keepdim=True) / n\n        x = (x - mu) / torch.sqrt(var + 1e-7)\n        x = x * mask\n    return x\n\n\ndef logits_b(x, mask):\n    return model(input_values=frontend_b(x, mask),\n                 attention_mask=mask.long()).logits\n\n\ndef ctc_per_sample(x, mask, tgt_pad, tgt_len):\n    \"\"\"Per-sample CTC loss -> lets us mask out finished samples.\"\"\"\n    lg = logits_b(x, mask)\n    logp = F.log_softmax(lg.float(), -1).transpose(0, 1)          # (L,B,V)\n    in_len = model._get_feat_extract_output_lengths(mask.sum(-1).long()).long()\n    in_len = in_len.clamp(max=logp.shape[0])\n    return F.ctc_loss(logp, tgt_pad, in_len, tgt_len,\n                      blank=BLANK_ID, zero_infinity=True, reduction=\"none\")\n\n\ndef decode_b(x, mask):\n    with torch.no_grad():\n        ids = torch.argmax(logits_b(x, mask), -1)\n    return [norm_bn(s) for s in processor.batch_decode(ids)]\n\n\ndef int16_rt(x):\n    return torch.clamp((x * 32767.0).round(), -32768.0, 32767.0) / 32767.0\n\ndef snr_db_b(clean, adv, mask):\n    n = (adv - clean) * mask\n    s = clean * mask\n    return (10 * torch.log10(s.pow(2).sum(-1) / (n.pow(2).sum(-1) + 1e-12))).cpu().numpy()","metadata":{},"outputs":[],"execution_count":null},{"id":"e16e052e-7edc-4238-9b55-91c14078d4f4","cell_type":"markdown","source":"## 6 · Batched PGD with per-sample budgets\n\nEvery utterance in the batch carries its **own** eps, its own success flag and its own\nbest-so-far delta. Finished samples are masked out of the loss so the optimiser stops spending\ncapacity on them, and their delta is frozen at the moment of success.","metadata":{}},{"id":"19dc4161-5bd0-4cfe-9602-1a7160e8bdb1","cell_type":"code","source":"def pgd_batch(x, mask, tgt_pad, tgt_len, targets, eps_vec,\n              steps, delta_init=None, lr=LR):\n    \"\"\"eps_vec: (B,) per-sample L-inf budgets.\n       Returns (success bool array, frozen deltas, steps_used).\"\"\"\n    B = x.shape[0]\n    eps_col = eps_vec.view(-1, 1)\n    delta = (torch.zeros_like(x) if delta_init is None\n             else torch.max(torch.min(delta_init.clone(), eps_col), -eps_col))\n    delta.requires_grad_(True)\n    opt = torch.optim.Adam([delta], lr=lr)\n    scaler = torch.amp.GradScaler(\"cuda\", enabled=USE_AMP)\n\n    active = torch.ones(B, dtype=torch.bool, device=DEVICE)\n    frozen = torch.zeros_like(x)\n    success = np.zeros(B, dtype=bool)\n    best_loss = torch.full((B,), float(\"inf\"), device=DEVICE)\n    stall = torch.zeros(B, dtype=torch.long, device=DEVICE)\n\n    for step in range(steps):\n        if not active.any():\n            break\n        x_adv = torch.clamp(x + delta, -1, 1) * mask\n        with torch.amp.autocast(\"cuda\", enabled=USE_AMP):\n            per = ctc_per_sample(x_adv, mask, tgt_pad, tgt_len)\n        loss = (per * active.float()).sum() / active.float().sum().clamp(min=1)\n\n        opt.zero_grad(); scaler.scale(loss).backward()\n        scaler.unscale_(opt)\n        if delta.grad is not None:\n            delta.grad[~active] = 0          # finished samples stop moving\n        scaler.step(opt); scaler.update()\n        with torch.no_grad():\n            delta.data = torch.max(torch.min(delta.data, eps_col), -eps_col)\n\n        with torch.no_grad():                # stall detection\n            improved = per < best_loss - 1e-3\n            best_loss = torch.minimum(best_loss, per.detach())\n            stall = torch.where(improved, torch.zeros_like(stall), stall + 1)\n            give_up = active & (stall > STALL_PATIENCE)\n            active = active & ~give_up\n\n        if (step + 1) % CHECK_EVERY == 0:\n            hyp = decode_b(int16_rt(torch.clamp(x + delta, -1, 1)) * mask, mask)\n            with torch.no_grad():\n                for i in range(B):\n                    if active[i] and hyp[i] == targets[i]:\n                        frozen[i] = delta[i].detach()\n                        success[i] = True\n                        active[i] = False\n    return success, frozen, step + 1","metadata":{},"outputs":[],"execution_count":null},{"id":"c86cac46-9412-4cd8-bc7a-0879c0e5e4d3","cell_type":"markdown","source":"## 7 · Per-sample geometric bisection, run in lockstep","metadata":{}},{"id":"d28ffe24-db7a-4a8c-902b-6b0175532ec4","cell_type":"code","source":"def eps50_batch(x, mask, tgt_pad, tgt_len, targets, verbose=True):\n    \"\"\"Per-sample geometric bisection for the smallest budget that still works.\n\n    ---------------------------------------------------------------------------\n    BUG FIXED 2026-09-22 (found in the rung-1 T5 run, 40% censoring).\n\n    The previous version ended with\n\n        ok, best, _ = pgd_batch(..., hi, CONFIRM_STEPS, delta_init=best)\n\n    which REASSIGNED `best` wholesale to pgd_batch's return value. That return\n    is a `frozen` buffer holding zeros for any sample that did not succeed on\n    that particular pass -- so one unlucky final pass wiped out a delta that had\n    already been verified at a larger budget. Those samples were then written\n    out as censored, with SNR 70-86 dB (the signature of delta == 0) even though\n    the ceiling probe had already proved them attackable.\n\n    In the rung-1 run that discarded ~6 of 25 real measurements. Worse, it does\n    so at a RATE THAT VARIES BY MODEL, so comparing median eps50 across rungs\n    would have been comparing differently-biased subsamples -- straight into\n    RQ1's headline.\n\n    Fix: keep a running `best_eps` / `best_delta` that is only ever updated on a\n    VERIFIED success. Nothing overwrites a confirmed result. Censoring now means\n    exactly one thing: the attack failed at every budget tried, ceiling included.\n\n    Side benefit: eps50 becomes a guaranteed UPPER BOUND on the true minimum.\n    Optimiser noise can only make it conservative, never wrong in an\n    unpredictable direction -- which is the property you want in a metric you\n    are about to compare across four models.\n    ---------------------------------------------------------------------------\n    \"\"\"\n    B = x.shape[0]\n    lo = torch.full((B,), EPS_LO, device=DEVICE)\n    hi = torch.full((B,), EPS_HI, device=DEVICE)\n    best_eps   = torch.full((B,), float(\"nan\"), device=DEVICE)   # verified only\n    best_delta = torch.zeros_like(x)\n\n    def record(okt, eps_vec, d):\n        \"\"\"Adopt a result only when it is both successful and better.\"\"\"\n        nonlocal best_eps, best_delta\n        better = okt & (torch.isnan(best_eps) | (eps_vec < best_eps))\n        best_eps   = torch.where(better, eps_vec, best_eps)\n        best_delta = torch.where(better.view(-1, 1), d, best_delta)\n\n    # Ceiling probe. Failure here is genuine right-censoring.\n    ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, hi, CONFIRM_STEPS)\n    okt = torch.tensor(ok, device=DEVICE)\n    record(okt, hi, d)\n    feasible = ok.copy()\n    if verbose:\n        print(f\"  ceiling probe eps={EPS_HI:.1e}: {feasible.sum()}/{B} feasible, \"\n              f\"{(~feasible).sum()} censored\")\n\n    for rnd in range(BISECT_ROUNDS):\n        mid = torch.sqrt(lo * hi)\n        ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, mid,\n                             PROBE_STEPS, delta_init=best_delta)\n        ok = ok & feasible\n        okt = torch.tensor(ok, device=DEVICE)\n        hi = torch.where(okt, mid, hi)\n        lo = torch.where(okt, lo, mid)\n        record(okt, mid, d)\n        if verbose:\n            print(f\"  round {rnd+1}/{BISECT_ROUNDS}: {ok.sum()}/{B} succeeded \"\n                  f\"| median eps {mid.median().item():.3e}\")\n\n    # Budget-consistency round: one more bisection at the FULL step budget, so\n    # the reported boundary is measured at the budget eps50 is defined at rather\n    # than at the cheaper probe budget. See RUNLOG for why this matters.\n    #\n    # Skipped in FIXED_EPS mode: there is no bracket to refine. The ceiling probe\n    # already ran at CONFIRM_STEPS, so the single reported budget is measured at\n    # the full budget and needs no second pass.\n    if BISECT_ROUNDS == 0:\n        eps       = best_eps.cpu().numpy()\n        confirmed = ~np.isnan(eps)\n        if verbose:\n            print(f\"  -> FIXED_EPS {EPS_HI:.1e}: {confirmed.sum()}/{B} succeeded, \"\n                  f\"{(~confirmed).sum()} resisted\")\n        return eps, np.ones(B), confirmed, best_delta\n\n    mid = torch.sqrt(lo * hi)\n    ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, mid,\n                         CONFIRM_STEPS, delta_init=best_delta)\n    ok = ok & feasible\n    okt = torch.tensor(ok, device=DEVICE)\n    record(okt, mid, d)\n    if verbose:\n        print(f\"  budget-consistency round @ {CONFIRM_STEPS} steps: {ok.sum()}/{B}\")\n\n    eps       = best_eps.cpu().numpy()\n    confirmed = ~np.isnan(eps)                 # censored == never succeeded\n    width     = (hi / lo).cpu().numpy()\n    if verbose:\n        print(f\"  -> {confirmed.sum()}/{B} measured, {(~confirmed).sum()} censored\")\n    return eps, width, confirmed, best_delta\n","metadata":{},"outputs":[],"execution_count":null},{"id":"4423d867-3ce1-438c-8e8d-cab3e6bc1622","cell_type":"markdown","source":"## 8 · Validate the batching before trusting it\n\nBatching is a substantial rewrite. If masked normalisation or the per-sample CTC lengths are\nwrong, results would be quietly biased rather than obviously broken. So: run two utterances\nbatched, then the same two alone, and compare.\n\n> **This costs about three extra eps50 searches (~1–2 GPU-h).** Run it once, for\n> `(RUNG=1, EDIT_TYPE=\"T5\")`, record the agreement in `RUNLOG.md`, then set\n> `VALIDATE_BATCHING = False` for the other seven runs. It is a one-off correctness check on\n> the batching code, and that code does not change between runs.","metadata":{}},{"id":"3ff8c416-2da7-4c81-b66c-10ae16229261","cell_type":"code","source":"if VALIDATE_BATCHING and len(ev) >= 2:\n    sub = ev.head(2)\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    print(\"batched (2 together):\")\n    t0 = time.time(); e_b, _, c_b, _ = eps50_batch(xb, mb, tpad, tl, tgts); tb = time.time()-t0\n\n    print(\"\\nsingle (one at a time):\")\n    e_s = []\n    t0 = time.time()\n    for i in range(2):\n        xs, ms = xb[i:i+1], mb[i:i+1]\n        e, _, c, _ = eps50_batch(xs, ms, tpad[i:i+1], tl[i:i+1], [tgts[i]])\n        e_s.append(e[0])\n    ts = time.time()-t0\n\n    print(f\"\\n{'':6s} {'batched':>12s} {'single':>12s}  {'rel diff':>9s}\")\n    for i in range(2):\n        a, b = e_b[i], e_s[i]\n        rd = \"n/a\" if (np.isnan(a) or np.isnan(b)) else f\"{100*abs(a-b)/b:8.1f}%\"\n        print(f\"  #{i}  {a:12.3e} {b:12.3e}  {rd:>9s}\")\n    print(f\"\\nwall-clock: batched {tb:.0f}s vs single {ts:.0f}s -> {ts/tb:.2f}x\")\n    print(\"Agreement within ~15% is fine (PGD is stochastic). Larger means a batching bug.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"63faa54a-2faf-4c0b-91a5-fb5863d67e70","cell_type":"markdown","source":"## 8b · Is eps50 converged at this step budget?\n","metadata":{}},{"id":"24282fa4-2fa1-486d-bc0f-78341e51e878","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# CONVERGENCE CHECK.\n#\n# eps50 is defined relative to CONFIRM_STEPS. If doubling the budget moves it\n# materially, the metric is still tracking optimiser convergence rather than\n# model robustness -- and since rungs may converge at different rates, that\n# would leak straight into the RQ1 comparison. Three utterances is enough to\n# detect it.\n# ---------------------------------------------------------------------------\nif len(ev) >= 3:\n    sub = ev.head(3)\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    _saved = CONFIRM_STEPS\n    print(f\"at CONFIRM_STEPS = {_saved}:\")\n    e1, _, _, _ = eps50_batch(xb, mb, tpad, tl, tgts, verbose=False)\n    CONFIRM_STEPS = _saved * 2\n    print(f\"at CONFIRM_STEPS = {CONFIRM_STEPS}:\")\n    e2, _, _, _ = eps50_batch(xb, mb, tpad, tl, tgts, verbose=False)\n    CONFIRM_STEPS = _saved\n\n    print(f\"\\n{'utt':>4s} {'@'+str(_saved):>12s} {'@'+str(_saved*2):>12s} {'shift':>9s}\")\n    shifts = []\n    for i in range(len(sub)):\n        a, b = e1[i], e2[i]\n        if np.isnan(a) or np.isnan(b):\n            print(f\"{i:>4d} {'censored':>12s}\"); continue\n        sh = abs(a - b) / a; shifts.append(sh)\n        print(f\"{i:>4d} {a:12.3e} {b:12.3e} {100*sh:8.1f}%\")\n    if shifts:\n        ms = float(np.median(shifts))\n        print(f\"\\nmedian shift on doubling the budget: {100*ms:.1f}%\")\n        if ms < 0.10:\n            print(\"Converged. eps50 is stable at this budget -- safe to compare rungs.\")\n        else:\n            print(\"NOT CONVERGED. Raise CONFIRM_STEPS and re-run, or state explicitly\")\n            print(\"that eps50 is budget-relative and hold the budget fixed across ALL\")\n            print(\"rungs so the comparison stays internally valid.\")\n        PROV[\"convergence_shift_pct\"] = round(100 * ms, 1)\n","metadata":{},"outputs":[],"execution_count":null},{"id":"7b3fcb41-d4a7-4637-a56c-9dc61d8cd753","cell_type":"markdown","source":"## 9 · The main sweep","metadata":{}},{"id":"36be8d0f-9472-48ef-8c0a-e4fae51bc090","cell_type":"code","source":"ev = ev.sort_values(\"duration\").reset_index(drop=True)   # length buckets = less padding\nOUT = f\"{WORK}/eps50_rung{RUNG}_{CHECKPOINT}_{EDIT_TYPE}.csv\"\n\n# Resume: if a previous session ran out of time, attach its output dataset as an\n# input and the utterances it already measured are carried over, not redone.\nresults = []\nfor p in glob.glob(f\"/kaggle/input/*/{os.path.basename(OUT)}\"):\n    prev = pd.read_csv(p)\n    results.extend(prev.to_dict(\"records\"))\n    print(f\"resuming: {len(prev)} utterances already measured in {os.path.basename(p)}\")\ndone = {r[\"utt_id\"] for r in results}\ntodo = ev[~ev.utt_id.isin(done)].reset_index(drop=True)\n\nbatches = [todo.iloc[i:i+BATCH_SIZE] for i in range(0, len(todo), BATCH_SIZE)]\nprint(f\"{len(todo)} utterances remaining -> {len(batches)} batches of up to {BATCH_SIZE}\\n\")\n\nDEADLINE = time.time() + SESSION_HOURS * 3600\nt_start = time.time()\n\nfor bi, sub in enumerate(batches):\n    if time.time() > DEADLINE:\n        print(f\"*** SESSION BUDGET ({SESSION_HOURS} h) reached. \"\n              f\"{sum(len(b) for b in batches[bi:])} utterances not yet measured. ***\\n\")\n        break\n\n    print(f\"=== batch {bi+1}/{len(batches)} ({len(sub)} utts, \"\n          f\"{sub.duration.min():.1f}-{sub.duration.max():.1f}s) ===\")\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    t0 = time.time()\n    eps, width, conf, best = eps50_batch(xb, mb, tpad, tl, tgts)\n    dt = time.time() - t0\n\n    x_adv = int16_rt(torch.clamp(xb + best, -1, 1)) * mb\n    snrs = snr_db_b(xb, x_adv, mb)\n    hyp = decode_b(x_adv, mb)\n\n    for i, (_, r) in enumerate(sub.iterrows()):\n        results.append(dict(\n            utt_id=r[\"utt_id\"], rung=RUNG, edit_type=EDIT_TYPE, checkpoint=CHECKPOINT,\n            duration=r[\"duration\"], source_text=r[\"text\"], target_text=r[\"target_text\"],\n            source_tokens=int(r[\"source_tokens\"]), target_tokens=int(r[\"target_tokens\"]),\n            eps50=float(eps[i]) if not np.isnan(eps[i]) else None,\n            # SNR/decode are meaningful only where an attack was verified.\n            # Writing them for censored rows produced the 70-86 dB values in\n            # the rung-1 run, which looked like data but were delta == 0.\n            censored=bool(np.isnan(eps[i])),\n            eps50_per_token=(float(eps[i]) / r[\"target_tokens\"]) if not np.isnan(eps[i]) else None,\n            bracket_width=float(width[i]), snr_db=float(snrs[i]),\n            decoded=(hyp[i] if not np.isnan(eps[i]) else None),\n            exact_match=(bool(hyp[i] == r[\"target_text\"])\n                         if not np.isnan(eps[i]) else False),\n            use_amp=USE_AMP, model_dir=MODEL_DIR,\n        ))\n\n    # Write after EVERY batch. The original only wrote at the very end, so a\n    # session killed on the last batch threw away hours of finished work.\n    pd.DataFrame(results).to_csv(OUT, index=False)\n\n    remaining = len(batches) - (bi + 1)\n    print(f\"  batch done in {dt/60:.1f} min | {(~np.isnan(eps)).sum()}/{len(sub)} uncensored \"\n          f\"| {len(results)} rows written\")\n    if remaining:\n        print(f\"  projected for the {remaining} remaining batches: \"\n              f\"{remaining*dt/3600:.1f} h \"\n              f\"({'fits' if time.time()+remaining*dt < DEADLINE else 'WILL NOT FIT this session'})\\n\")\n    else:\n        print()\n\nprint(f\"TOTAL {(time.time()-t_start)/60:.1f} min this session | {len(results)}/{len(ev)} measured\")","metadata":{},"outputs":[],"execution_count":null},{"id":"3094ec17-5270-4100-a0b7-39ce062e9d2c","cell_type":"markdown","source":"## 10 · Results and save","metadata":{}},{"id":"b60bacc7-3530-4be1-9824-5d0af4a33905","cell_type":"code","source":"CK_TAG = \"final\" if CHECKPOINT == \"final\" else \"ms\" + CHECKPOINT.split(\"-\")[1]\nif FIXED_EPS is not None:\n    CK_TAG += \"-fixedeps\"        # keep trajectory runs separate from eps50 runs\n# Every checkpoint of a rung gets its OWN dataset. Before this, 100h final and\n# all three 100h milestones were told to publish as \"eps50-rung100-T5\", so\n# \"New Version\" on that dataset silently replaced one checkpoint's results\n# with another's.\ndf = pd.DataFrame(results)\nok = df[~df.censored] if len(df) else df\ncomplete = len(df) >= len(ev)\n\nprint(f\"=== RUNG {RUNG}h | {EDIT_TYPE}\"\n      + (f\" | FIXED_EPS {FIXED_EPS:.1e} ===\" if FIXED_EPS is not None else \" ===\"))\nif FIXED_EPS is not None:\n    print(\"  FIXED_EPS mode: every success is recorded at exactly the fixed budget, so\")\n    print(\"  eps50, its IQR and its bootstrap CI below are DEGENERATE and carry no\")\n    print(\"  information. Use ASR@eps and the per-utterance `censored` column only.\")\nprint(f\"utterances      : {len(df)} of {len(ev)}\")\nprint(f\"censored        : {df.censored.sum()} ({100*df.censored.mean():.0f}%)  <- report this\")\nif len(ok):\n    print(f\"median eps50    : {ok.eps50.median():.3e}  ({ok.eps50.median()*32767:.1f} int16 units)\")\n    print(f\"IQR             : {ok.eps50.quantile(.25):.3e} - {ok.eps50.quantile(.75):.3e}\")\n    print(f\"per target token: {ok.eps50_per_token.median():.3e}   <- length-normalised\")\n    print(f\"median SNR      : {ok.snr_db.median():.1f} dB\")\n    print(f\"exact match     : {ok.exact_match.sum()}/{len(ok)}\")\n\n    # A bracket width near 1.0 with eps50 at the floor means the search hit\n    # EPS_LO: LEFT-censored, and just as biasing as the right-censored kind.\n    at_floor = int(((ok.eps50 <= EPS_LO * 1.05)).sum())\n    if at_floor:\n        print(f\"WARNING: {at_floor} utterances sit at the EPS_LO floor ({EPS_LO:.0e}) \"\n              \"-- left-censored. Lower EPS_LO and re-run this rung.\")\n\n    b = [ok.eps50.sample(len(ok), replace=True).median() for _ in range(2000)]\n    print(f\"bootstrap 95% CI: [{np.percentile(b,2.5):.3e}, {np.percentile(b,97.5):.3e}]\")\n\ndf.to_csv(f\"{WORK}/eps50_rung{RUNG}_{CHECKPOINT}_{EDIT_TYPE}.csv\", index=False)\nPROV.update(n_utts=len(df), n_expected=len(ev), complete=bool(complete),\n            n_censored=int(df.censored.sum()) if len(df) else 0,\n            median_eps50=float(ok.eps50.median()) if len(ok) else None,\n            median_eps50_per_token=float(ok.eps50_per_token.median()) if len(ok) else None,\n            median_snr=float(ok.snr_db.median()) if len(ok) else None,\n            minutes=round((time.time()-t_start)/60, 1))\njson.dump(PROV, open(f\"{WORK}/_provenance.json\", \"w\"), indent=2)\n\nKAGGLE_USER = os.environ.get(\"KAGGLE_USERNAME\") or \"YOURUSERNAME\"\njson.dump({\"title\": f\"eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}\",\n           \"id\": f\"{KAGGLE_USER}/eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}\",\n           \"licenses\": [{\"name\": \"CC0-1.0\"}]},\n          open(f\"{WORK}/dataset-metadata.json\", \"w\"), indent=2)\n\nif not complete:\n    print(\"\\n\" + \"=\" * 70)\n    print(f\"  PARTIAL: {len(df)}/{len(ev)} utterances measured. Nothing is lost.\")\n    print(f\"    1. Output panel -> Create Dataset (or New Version) \"\n          f\"-> eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}\")\n    print( \"    2. Attach that dataset back as an INPUT to this notebook\")\n    print( \"    3. Re-commit. It skips what is already done.\")\n    print(\"=\" * 70)\nelse:\n    print(f\"\\nRUNG {RUNG}h / {EDIT_TYPE} COMPLETE -- publish as \"\n          f\"eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()} and record the median in RUNLOG.md.\")\n\nprint(\"\\n\" + json.dumps(PROV, indent=2))\ndf.head(10)\n\n# ---------------------------------------------------------------------------\n# CO-PRIMARY METRIC: attack success rate at a FIXED budget.\n#\n# Median eps50 is computed only over UNCENSORED utterances, so if the censoring\n# rate differs between rungs -- and it will -- the medians describe different\n# subpopulations and are not directly comparable. That is a censoring bias\n# sitting directly under RQ1's headline.\n#\n# ASR@ceiling has no such problem: every utterance contributes, success or\n# failure, so it is comparable across rungs by construction.\n#\n# The two measure different things and should both be reported:\n#   ASR@ceiling -> can this model be STEERED to an exact target at all?\n#   median eps50 -> when it can, how much perturbation does it take?\n#\n# Expect them to move in opposite directions. A better model is usually easier\n# to aim (lower censoring) but harder to push (higher eps50). Reporting only\n# one of them would tell half the story, and the flattering half depends on\n# which one you pick.\n# ---------------------------------------------------------------------------\nASR_AT_CEILING = float((~df.censored).mean())\nprint(f\"\\nASR@eps={EPS_HI:.0e}   : {ASR_AT_CEILING:.1%}  ({(~df.censored).sum()}/{len(df)})\")\nprint(f\"  <- censoring-free, uses all {len(df)} utterances, comparable across rungs\")\nprint(f\"median eps50 is over the {(~df.censored).sum()} steerable utterances only\")\n\n# ---------------------------------------------------------------------------\n# LEFT-censoring: the mirror image of the block above, and the one this notebook\n# had no check for. If eps50 lands at the bottom of the bracket, the attack\n# succeeded at the smallest budget searched, so the value is an UPPER BOUND and\n# the true eps50 is somewhere below it. Reporting it as a point estimate\n# understates how cheap the attack is -- in the direction that flatters RQ2.\n# ---------------------------------------------------------------------------\nfloor_hits = int((ok.eps50 <= EPS_LO * 1.15).sum()) if len(ok) else 0\nPROV[\"left_censored\"] = floor_hits\nif floor_hits:\n    print(f\"\\n*** LEFT-CENSORED: {floor_hits}/{len(ok)} eps50 values sit at the \"\n          f\"bracket floor ({EPS_LO:.1e}) ***\")\n    print(\"    These are upper bounds, not measurements. Lower EPS_LO by a decade\")\n    print(\"    and re-run, or report them as censored-below and say so in the paper.\")\n\nif len(ok):\n    b = [ok.eps50.sample(len(ok), replace=True).median() for _ in range(2000)]\n    lo_ci, hi_ci = np.percentile(b, 2.5), np.percentile(b, 97.5)\n    print(f\"\\neps50 bootstrap 95% CI spans {hi_ci/lo_ci:.2f}x\")\n    if hi_ci / lo_ci > 1.8:\n        print(\"  WIDE. Between-rung differences must exceed this to be credible.\")\n        print(\"  If the ladder effect is smaller than the CI, raise CONFIRM_STEPS or\")\n        print(\"  run each condition twice -- but only if the effect turns out marginal.\")\n\nPROV[\"asr_at_ceiling\"] = round(ASR_AT_CEILING, 4)\nPROV[\"eps50_ci_ratio\"] = round(float(hi_ci / lo_ci), 3) if len(ok) else None\n","metadata":{},"outputs":[],"execution_count":null},{"id":"420ba1e9-20fc-4960-9d58-2392fb5db494","cell_type":"markdown","source":"---\n\n## Record this in RUNLOG.md, then run the next combination\n\n| Rung | T1 median eps50 | T1 censored | T5 median eps50 | T5 censored |\n|---|---|---|---|---|\n| 1 h | | | | |\n| 5 h | | | | |\n| 25 h | | | | |\n| 100 h | | | | |\n\n### What you are looking for\n\n**RQ1:** eps50 should *rise* with training hours — more data, more perturbation needed. That is\nthe headline. It only counts after NB-02's iso-WER adjustment, since clean WER is a covariate.\n\n**RQ2:** T1 should cost less than T5 at matched target length. Compare the\n`eps50_per_token` column, not the raw one.\n\n### Watch the censoring rate\n\nIf the 1 h rung shows **high censoring**, that is the paradox the plan predicted: a model whose\noutput is already unstable can be easy to *disrupt* yet hard to steer to a *specific* target.\nThat is a finding, not a failure. Report censoring alongside eps50 and use Kaplan–Meier rather\nthan silently dropping those rows.\n\n### Next\n**NB-06 — the LM-weight sweep.** Cheapest experiment in the thesis and the one that turns the\ncorrelation into a mechanism. Prioritise it over everything except the ladder itself.","metadata":{}}]}