{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"accelerator":"GPU","task_b_adaptation":{"source_notebook_sha256":"a673cbf4b1679bd227f47b2771f31180905e0dfbeec5731ea6b8134dd2a504d4","adapter_sha256":"abe2bbe01b0599eae447d8db86a9a09d57c365ff2f1200fd7a2ddce6f0a60a3d","attack_cells_unchanged":[11,13,15,17,19],"validation_status":"CPU/static only; actual Kaggle data and GPU execution pending"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 1 · Configuration","metadata":{}},{"cell_type":"markdown","source":"# Task B - NB-05, 100 recordings, six locked conditions\n\nPrepare now; run GPU experiments after Task A is complete, as the coworker requires.\nThis is a separate adaptation of the supplied NB-05. Its front-end, PGD, bisection,\nbatching diagnostic, and convergence diagnostic code are preserved.\n\nChange **RUN_ID only** for a fresh experiment. All runs use T5, n=100, full bisection,\nAMP, seed 1337, batch size 8, and the original attack budgets.\n\n| RUN_ID | Rung | Checkpoint | Checkpoint input | Separate output dataset |\n|---|---|---|---|---|\n| 1 | 1 | final | ckpt-xlsr-1h | eps50-rung1-final-t5-n100 |\n| 2 | 5 | final | ckpt-xlsr-5h | eps50-rung5-final-t5-n100 |\n| 3 | 25 | final | ckpt-xlsr-25h | eps50-rung25-final-t5-n100 |\n| 4 | 100 | milestone-1000 | ckpt-xlsr-100h | eps50-rung100-ms1000-t5-n100 |\n| 5 | 100 | milestone-2000 | ckpt-xlsr-100h | eps50-rung100-ms2000-t5-n100 |\n| 6 | 100 | milestone-3500 | ckpt-xlsr-100h | eps50-rung100-ms3500-t5-n100 |\n\n**Always attach:** bn-asr-splits, bn-attack-eval-100, bengaliai-speech competition audio,\nand the checkpoint input above. Do not attach bn-attack-targets or substitute ckpt-xlsr-1h-traj.\nThe n100 CSV is checked against the SHA-256 reported when it was published.\n\nThe configuration and file-preflight cells can run on CPU. Stop before GPU initialization\nif only preparing. For the full run use the supplied NB-05 setup: GPU T4 x2, Internet On,\nand Save Version / Save & Run All. The attack code exposes only one GPU.\nStart with a fresh session/output directory for each condition. Run one condition at a time,\npublish its separate n100 output, and append its RUNLOG entry before the next condition.\n\nFor a partial-run recovery only, set RESUME_DATASET to that condition's exact published\ndataset name. Resume verifies hashes, settings, checkpoint weights, targets, and result rows.\nOld n25 datasets are never reused. Keep the same RUN_ID. A recovered run is not bitwise\nequivalent to an uninterrupted run; record the restart in RUNLOG.\n\nThe 100 targets and their SHA-256 are exported. Compare target hashes across all six runs.\nThe original 25 keep their order before target generation. Historical target equivalence\nstill requires comparison with an actual historical T5 output; that data is not bundled here.\n\nDo not change a condition's attack budget in response to a convergence/floor warning.\nRecord the diagnostic and discuss it with the lead author. The original budget-relative\nmeasurement is retained for comparability. Censored rows remain in the result.\n","metadata":{}},{"cell_type":"code","source":"RUN_ID = 6 # EDIT THIS ONLY for a fresh condition: 1, 2, 3, 4, 5, 6.\nRESUME_DATASET = None  # Recovery only: exact published partial n100 dataset name.\n\n_CONDITIONS = {\n    1: (1, \"final\"), 2: (5, \"final\"), 3: (25, \"final\"),\n    4: (100, \"milestone-1000\"), 5: (100, \"milestone-2000\"),\n    6: (100, \"milestone-3500\"),\n}\nassert type(RUN_ID) is int and RUN_ID in _CONDITIONS, \"RUN_ID must be 1 through 6.\"\nRUNG, CHECKPOINT = _CONDITIONS[RUN_ID]\nEDIT_TYPE = \"T5\"\nN_EVAL = 100\nEVAL_SET = \"bn-attack-eval-100\"\nFIXED_EPS = None\nCKPT_DATASET = None\nBATCH_SIZE = 8\nUSE_AMP = True\nSESSION_HOURS = 8.0\nEPS_LO, EPS_HI = 1e-4, 2e-3\nBISECT_ROUNDS = 5\nPROBE_STEPS = 400\nCONFIRM_STEPS = 800\nSTALL_PATIENCE = 100\nCHECK_EVERY = 50\nLR = 1e-4\nVALIDATE_BATCHING = False\nEXPECTED_EVAL_SHA256 = \"93e69343d5d66727b3d2dd343dd031b7a9a425104afb45011cdd50b4a94ad7cf\"\nCK_TAG = \"final\" if CHECKPOINT == \"final\" else \"ms\" + CHECKPOINT.split(\"-\")[1]\nOUTPUT_DATASET = f\"eps50-rung{RUNG}-{CK_TAG}-t5-n100\"\nassert RESUME_DATASET is None or RESUME_DATASET == OUTPUT_DATASET, (\n    \"Resume must use this exact condition's published n100 dataset.\"\n)\n\ndef assert_task_b_settings():\n    actual = (RUNG, CHECKPOINT, EDIT_TYPE, N_EVAL, EVAL_SET, FIXED_EPS, CKPT_DATASET,\n              BATCH_SIZE, USE_AMP, SESSION_HOURS, EPS_LO, EPS_HI, BISECT_ROUNDS,\n              PROBE_STEPS, CONFIRM_STEPS, STALL_PATIENCE, CHECK_EVERY, LR, VALIDATE_BATCHING)\n    expected = (*_CONDITIONS[RUN_ID], \"T5\", 100, \"bn-attack-eval-100\", None, None,\n                8, True, 8.0, 1e-4, 2e-3, 5, 400, 800, 100, 50, 1e-4, False)\n    assert actual == expected, \"Task B configuration drift. Restore the supplied settings.\"\n\nassert_task_b_settings()\nprint(f\"RUN_ID={RUN_ID}: rung {RUNG}, {CHECKPOINT}, T5, n100, full bisection\")\nprint(\"Output dataset:\", OUTPUT_DATASET)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-28T04:07:12.541606Z","iopub.execute_input":"2026-09-28T04:07:12.541874Z","iopub.status.idle":"2026-09-28T04:07:12.558168Z","shell.execute_reply.started":"2026-09-28T04:07:12.541838Z","shell.execute_reply":"2026-09-28T04:07:12.557203Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## CPU file preflight\nRun through this cell to prepare without running attacks.\n","metadata":{}},{"cell_type":"code","source":"\"\"\"Embedded into the Task B notebook; runs before importing torch or loading weights.\"\"\"\nimport glob\nimport hashlib\nimport json\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\n\ndef sha256_file(path):\n    return hashlib.sha256(Path(path).read_bytes()).hexdigest()\n\n\ndef one_file(root, name):\n    matches = sorted(Path(root).rglob(name))\n    assert len(matches) == 1, f\"Need exactly one {name} in {root}; found {matches}\"\n    return matches[0]\n\n\ndef task_b_preflight(input_root, expected_sha, rung, checkpoint):\n    roots = sorted({Path(p) for depth in (1, 2, 3)\n                    for p in glob.glob(str(Path(input_root).joinpath(*(['*'] * depth))))\n                    if Path(p).is_dir()})\n\n    def dataset(name):\n        found = [p for p in roots if p.name == name]\n        assert len(found) == 1, f\"Attach exactly one {name}; found {found}\"\n        return found[0]\n\n    split_root = dataset('bn-asr-splits')\n    eval_root = dataset('bn-attack-eval-100')\n    audio_root = dataset('bengaliai-speech')\n    checkpoint_root = dataset(f'ckpt-xlsr-{rung}h')\n    model_dir = (checkpoint_root if (checkpoint_root / 'config.json').is_file()\n                 else checkpoint_root / 'final') if checkpoint == 'final' else (\n        checkpoint_root / 'milestones' / f\"step-{checkpoint.split('-')[1]}\")\n    assert (model_dir / 'config.json').is_file(), f\"Checkpoint missing: {model_dir}\"\n\n    eval_csv = one_file(eval_root, 'attack_eval_100.csv')\n    assert sha256_file(eval_csv) == expected_sha, (\n        'The n100 CSV differs from the published selection. Check the attached dataset version.'\n    )\n    selection_prov_path = one_file(eval_root, 'attack_eval_100_provenance.json')\n    selection_prov = json.loads(selection_prov_path.read_text(encoding='utf-8'))\n    assert selection_prov['output_sha256'] == expected_sha, 'Selection provenance hash mismatch.'\n    assert selection_prov['seed'] == 1337, 'Wrong selection seed.'\n\n    old_path = one_file(split_root, 'attack_eval_25.csv')\n    splits = old_path.parent\n    test_path = splits / 'test.csv'\n    rung_paths = sorted(splits.glob('rung_*.csv'))\n    assert {f'rung_{n}h.csv' for n in (1, 5, 25, 100)} <= {p.name for p in rung_paths}, (\n        'Missing training-rung manifests.'\n    )\n    source_paths = [old_path, test_path, *rung_paths]\n    source_hashes = {p.name: sha256_file(p) for p in source_paths}\n    assert source_hashes == selection_prov['source_file_sha256'], (\n        'Frozen split files differ from the versions used to select the n100 set.'\n    )\n\n    frozen = pd.read_csv(eval_csv, dtype={'utt_id': str})\n    old = pd.read_csv(old_path, dtype={'utt_id': str})\n    test = pd.read_csv(test_path, dtype={'utt_id': str})\n    ids = frozen.utt_id.tolist()\n    assert len(old) == 25 and old.utt_id.is_unique\n    assert len(frozen) == 100 and frozen.utt_id.notna().all() and len(set(ids)) == 100\n    assert ids[:25] == old.utt_id.tolist(), 'Original 25 IDs/order changed.'\n    pd.testing.assert_frame_equal(frozen.iloc[:25].reset_index(drop=True), old,\n                                  check_dtype=False, check_exact=False)\n    assert test.utt_id.notna().all() and test.utt_id.is_unique\n    assert set(ids) <= set(test.utt_id), 'A selected ID is outside test.csv.'\n    for p in rung_paths:\n        train = pd.read_csv(p, usecols=['utt_id'], dtype={'utt_id': str})\n        assert train.utt_id.notna().all(), f'Missing training ID: {p.name}'\n        assert not (set(ids) & set(train.utt_id)), f'Training overlap: {p.name}'\n    assert np.isfinite(frozen.duration).all() and frozen.duration.between(0, 10, inclusive='right').all()\n    assert frozen['text'].notna().all() and frozen['text'].str.strip().ne('').all()\n\n    # Resolve only within the attached competition dataset; never search every audio file.\n    resolved = []\n    for raw in frozen.path:\n        original = Path(raw)\n        if original.is_file() and original.resolve().is_relative_to(audio_root.resolve()):\n            actual = original\n        else:\n            actual = audio_root / original.parent.name / original.name\n        assert actual.is_file(), f'Audio missing: {actual}. Check the competition input.'\n        resolved.append(str(actual))\n    frame = frozen.copy()\n    frame['path'] = resolved\n    summary = pd.DataFrame({\n        'original_25': frozen.iloc[:25].duration.agg(['min', 'median', 'max']),\n        'additional_75': frozen.iloc[25:].duration.agg(['min', 'median', 'max']),\n    })\n    print('PREFLIGHT PASS: published CSV hash, original 25, 100 unique test IDs, no training overlap.')\n    print('PREFLIGHT PASS: 100 audio files present; checkpoint:', model_dir)\n    print(summary.to_string())\n    source_version = selection_prov.get('source_dataset_version')\n    if not source_version or str(source_version).startswith('Record from'):\n        print('AUDIT NOTE: source dataset version was not filled in; record it in RUNLOG. File hashes verified.')\n    return dict(\n        frame=frame, ids=ids, inputs=[str(p) for p in roots],\n        splits=str(splits), eval_csv=str(eval_csv), model_dir=str(model_dir),\n        checkpoint_root=str(checkpoint_root), audio_root=str(audio_root),\n        provenance={\n            'attack_eval_sha256': expected_sha,\n            'selection_provenance_sha256': sha256_file(selection_prov_path),\n            'source_file_sha256': source_hashes,\n            'source_dataset_version': source_version,\n            'selection_provenance': selection_prov,\n            'expected_ids': ids,\n        },\n    )\n\nassert_task_b_settings()\nPREFLIGHT = task_b_preflight('/kaggle/input', EXPECTED_EVAL_SHA256, RUNG, CHECKPOINT)\nprint('CPU file preflight finished. GPU initialization is the next section.')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-28T04:08:58.618341Z","iopub.execute_input":"2026-09-28T04:08:58.618683Z","iopub.status.idle":"2026-09-28T04:09:00.968194Z","shell.execute_reply.started":"2026-09-28T04:08:58.618653Z","shell.execute_reply":"2026-09-28T04:09:00.967222Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2 · Boilerplate","metadata":{}},{"cell_type":"code","source":"import os\nos.environ[\"CUDA_VISIBLE_DEVICES\"] = \"0\"   # attacks optimise the INPUT: never DataParallel\nimport sys, json, time, random, re, glob, unicodedata\nimport numpy as np, pandas as pd, torch\nimport torch.nn.functional as F\ntorch.backends.cudnn.benchmark = True\n\nSEED = 1337\nrandom.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)\nWORK, TEMP, DEVICE = \"/kaggle/working\", \"/kaggle/temp\", \"cuda\"\nassert torch.cuda.is_available(), \"Settings -> Accelerator -> GPU T4 x2\"\n\nPROV = {\"notebook\": \"NB-05-task-b-n100-locked\", \"rung\": RUNG, \"edit_type\": EDIT_TYPE,\n        \"n_eval\": N_EVAL, \"use_amp\": USE_AMP, \"seed\": SEED,\n        \"bracket\": [EPS_LO, EPS_HI], \"bisect_rounds\": BISECT_ROUNDS,\n        \"gpu\": torch.cuda.get_device_name(0), \"torch\": torch.__version__,\n        \"timestamp\": time.strftime(\"%Y-%m-%d %H:%M:%S\")}\nprint(json.dumps(PROV, indent=2))\nassert_task_b_settings()\nassert SEED == 1337\nNOTEBOOK_STARTED = time.time()\nPath(WORK).mkdir(parents=True, exist_ok=True)\nassert not list(Path(WORK).glob('eps50*.csv')), \"Use a fresh session for this condition.\"\nassert not Path(WORK, '_provenance.json').exists(), \"Use a fresh session; publish existing work first.\"\nPROV.update(PREFLIGHT['provenance'])\nPROV.update(source_notebook_sha256='a673cbf4b1679bd227f47b2771f31180905e0dfbeec5731ea6b8134dd2a504d4', adapter_sha256='abe2bbe01b0599eae447d8db86a9a09d57c365ff2f1200fd7a2ddce6f0a60a3d',\n            checkpoint=CHECKPOINT, batch_size=BATCH_SIZE, probe_steps=PROBE_STEPS,\n            confirm_steps=CONFIRM_STEPS, stall_patience=STALL_PATIENCE,\n            check_every=CHECK_EVERY, lr=LR, pandas=pd.__version__,\n            output_dataset=OUTPUT_DATASET, resume_dataset=RESUME_DATASET)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# See NB-01 for the full explanation: pinning an old `accelerate`/`datasets`\n# silently downgrades numpy mid-session and breaks pandas much later with\n# \"No module named 'numpy.rec'\". Pin numpy to what is already loaded.\n# `datasets` is not imported by this notebook, so it is not installed.\nimport numpy, subprocess, sys\nNP_BEFORE = numpy.__version__\n\n!pip install -q transformers==4.44.2 librosa soundfile jiwer \"numpy=={NP_BEFORE}\" 2>&1 | tail -2\n\nNP_DISK = subprocess.run([sys.executable, \"-c\", \"import numpy; print(numpy.__version__)\"],\n                         capture_output=True, text=True).stdout.strip()\nassert NP_BEFORE == NP_DISK, f\"numpy moved {NP_BEFORE} -> {NP_DISK}; relax the pin that forced it.\"\nprint(f\"numpy {NP_BEFORE} stable\")\n\nfrom transformers import Wav2Vec2ForCTC, Wav2Vec2Processor\nimport transformers, librosa\nprint(\"transformers\", transformers.__version__)\nPROV[\"transformers\"] = transformers.__version__\nPROV[\"numpy\"] = NP_BEFORE\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3 · Load the ladder checkpoint and the frozen attack set","metadata":{}},{"cell_type":"code","source":"inputs = PREFLIGHT[\"inputs\"]\n# Match the rung EXACTLY. The old test was `f\"{RUNG}h\" in p`, and \"5h\" is a\n# substring of \"ckpt-xlsr-25h\" -- with both datasets attached, RUNG=5 could\n# silently attack the 25 h model and corrupt the headline figure.\n_DS = CKPT_DATASET or f\"ckpt-xlsr-{RUNG}h\"\ncands = [p for p in inputs if re.search(rf\"{re.escape(_DS)}$\", os.path.basename(p))]\nassert len(cands) == 1, (\n    f\"Need exactly one input matching '{_DS}', found {len(cands)}: \"\n    f\"{[os.path.basename(c) for c in cands]}\\n\"\n    f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\nCKPT = cands[0]\nif CHECKPOINT == \"final\":\n    MODEL_DIR = CKPT if os.path.exists(f\"{CKPT}/config.json\") else f\"{CKPT}/final\"\nelse:\n    MODEL_DIR = f\"{CKPT}/milestones/step-{CHECKPOINT.split('-')[1]}\"\n    assert os.path.isdir(MODEL_DIR), (\n        f\"{MODEL_DIR} missing. Milestones live at steps 1000/2000/3500. They exist \"\n        f\"for rungs >= 25h in ckpt-xlsr-{{25,100}}h, and for the 1h rung ONLY in \"\n        f\"ckpt-xlsr-5h-traj (the 2026-09-26 TRAJ_RUN retrain) -- the original \"\n        f\"ckpt-xlsr-5h has no milestones/ dir. If you want 1h milestones, set \"\n        f\"CKPT_DATASET = 'ckpt-xlsr-1h-traj'.\")\nassert os.path.exists(f\"{MODEL_DIR}/config.json\"), \\\n    f\"no config.json under {MODEL_DIR} -- did NB-01 finish and publish this rung?\"\nprint(\"checkpoint:\", MODEL_DIR)\n\nprocessor = Wav2Vec2Processor.from_pretrained(MODEL_DIR)\nmodel = Wav2Vec2ForCTC.from_pretrained(MODEL_DIR).float().to(DEVICE).eval()\nfor p in model.parameters(): p.requires_grad_(False)\nmodel.config.apply_spec_augment = False      # no masking noise during attacks\n\nDO_NORMALIZE = bool(getattr(processor.feature_extractor, \"do_normalize\", True))\nBLANK_ID = model.config.pad_token_id\nprint(f\"do_normalize {DO_NORMALIZE} | blank {BLANK_ID} | vocab {model.config.vocab_size}\")\n\nassert Path(MODEL_DIR) == Path(PREFLIGHT['model_dir'])\nSPLITS = PREFLIGHT['splits']\nEV_CSV = PREFLIGHT['eval_csv']\nev = PREFLIGHT['frame'].copy()\nEXPECTED_IDS = PREFLIGHT['ids']\nAUDIO_ROOT = PREFLIGHT['audio_root']\nassert EVAL_SET == \"bn-attack-eval-100\" and EDIT_TYPE == \"T5\" and N_EVAL == 100\nPROV.update(attack_eval_csv=\"attack_eval_100.csv\", fixed_eps=FIXED_EPS,\n            ckpt_dataset=_DS, bisect_rounds=BISECT_ROUNDS,\n            model_dir=MODEL_DIR, audio_root=AUDIO_ROOT)\nprint(f\"eval set: {EV_CSV}; {len(ev)} utterances; all audio present\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4 - Build the original T5 length-matched targets\n\nAutomatic targets from frozen test.csv are the intended RQ1 procedure. Preserve CSV row\norder until target generation is complete. The new mode deliberately ignores the T1 corpus.\n","metadata":{}},{"cell_type":"code","source":"ZW = dict.fromkeys(map(ord, \"‌‍\"), None)\ndef norm_bn(s):\n    return re.sub(r\"\\s+\", \" \", unicodedata.normalize(\"NFC\", str(s)).translate(ZW)).strip()\ndef ntok(s): return len(processor.tokenizer(norm_bn(s)).input_ids)\n\n# Use the resolved `inputs` list from the cell above, NOT a flat glob. Your own\n# datasets mount at /kaggle/input/datasets/<user>/<name>, three levels deep, so\n# glob(\"/kaggle/input/*\") sees only the wrapper dirs. That is the same failure\n# NB-01 hit, and on 2026-09-25 it silently hid an attached bn-attack-targets:\n# T5 then auto-generated targets on the wrong utterance set and still \"succeeded\".\n# The TARGET source is tied to EVAL_SET so the two can never diverge. Under\n# \"t5_frozen\" the NB-03 corpus is not consulted even when attached: its utt_ids\n# are disjoint from attack_eval_25.csv, so a merge would inner-join to zero rows.\nTGT_CSV = None\nif EVAL_SET == \"nb03\":\n    TGT_CSV = next((f\"{p}/bn_attack_targets.csv\" for p in inputs\n                    if os.path.exists(f\"{p}/bn_attack_targets.csv\")), None)\n    assert TGT_CSV, (\n        \"EVAL_SET='nb03' but bn_attack_targets.csv not found. Attach \"\n        \"bn-attack-targets (run NB-03 first).\\n\"\n        f\"Visible inputs: {[os.path.basename(p) for p in inputs]}\")\n\nif TGT_CSV:\n    tg = pd.read_csv(TGT_CSV)\n    tg = tg[tg.edit_type == EDIT_TYPE]\n    ev = ev.merge(tg[[\"source_utterance_id\", \"target_text\"]],\n                  left_on=\"utt_id\", right_on=\"source_utterance_id\", how=\"inner\")\n    print(f\"using NB-03 hand-built corpus: {len(ev)} pairs\")\nelse:\n    print(\"NB-03 corpus not attached -> generating targets automatically.\")\n    print(\"Task B uses the original automatic T5 targets.\\n\")\n    pool = pd.read_csv(f\"{SPLITS}/test.csv\")\n    pool[\"ntok\"] = pool[\"text\"].map(ntok)\n    ev[\"src_ntok\"] = ev[\"text\"].map(ntok)\n\n    rows = []\n    rng = np.random.RandomState(SEED)\n    for _, r in ev.iterrows():\n        srct = norm_bn(r[\"text\"])\n        if EDIT_TYPE == \"T1\":\n            # STANDALONE particle only. A substring test is wrong: \"না\" occurs\n            # inside many common Bengali words -- নাম (name), নারী (woman),\n            # জানা (to know) -- and str.replace() would carve syllables out of\n            # them, producing a target that is not a negation inversion at all.\n            # Remove the LAST standalone \"না\", which is the post-verbal negation.\n            toks = srct.split()\n            if \"না\" not in toks: continue\n            j = len(toks) - 1 - toks[::-1].index(\"না\")\n            tgt = norm_bn(\" \".join(toks[:j] + toks[j+1:]))\n            if tgt == srct or not tgt: continue\n        else:\n            cand = pool[(pool.utt_id != r[\"utt_id\"]) &\n                        (pool.ntok.between(r[\"src_ntok\"] - 2, r[\"src_ntok\"] + 2))]\n            if len(cand) == 0: continue\n            tgt = norm_bn(cand.sample(1, random_state=int(rng.randint(1e6))).iloc[0][\"text\"])\n        rows.append({**r.to_dict(), \"target_text\": tgt})\n    ev = pd.DataFrame(rows)\n    if EDIT_TYPE == \"T1\":\n        # NB-00 froze the attack set using a SUBSTRING test for \"না\", so the\n        # advertised \"15 with না\" overcounts utterances that actually carry the\n        # standalone negation particle. Find out before spending GPU.\n        print(f\"\\nT1: {len(ev)} of {N_EVAL} utterances carry a standalone না\")\n        assert len(ev) >= 10, (\n            f\"only {len(ev)} T1-attackable utterances -- too few for RQ2. \"\n            f\"Either build the NB-03 corpus, or freeze an additional T1-specific \"\n            f\"attack subset from the test split (without touching bn-asr-splits).\")\n\nev[\"target_text\"] = ev[\"target_text\"].map(norm_bn)\nev[\"target_tokens\"] = ev[\"target_text\"].map(ntok)\nev[\"source_tokens\"] = ev[\"text\"].map(ntok)\nev = ev.reset_index(drop=True)\nPROV[\"n_attackable\"] = int(len(ev))\nPROV[\"n_eval_frozen\"] = int(N_EVAL)\nPROV[\"mean_target_tokens\"] = float(ev.target_tokens.mean())\nPROV[\"mean_source_tokens\"] = float(ev.source_tokens.mean())\nprint(f\"{EDIT_TYPE}: {len(ev)} attackable utterances\")\nprint(f\"target tokens  mean {ev.target_tokens.mean():.1f}  vs source {ev.source_tokens.mean():.1f}\")\nprint(\"  (these should be close for T5; T1 is by construction slightly shorter)\")\nev[[\"utt_id\", \"source_tokens\", \"target_tokens\", \"text\", \"target_text\"]].head()\nassert len(ev) == N_EVAL and ev.utt_id.tolist() == EXPECTED_IDS, (\n    \"Target generation lost or reordered recordings. Stop before attacks; do not resample silently.\"\n)\nassert ev.target_text.str.strip().ne('').all() and ev.target_tokens.gt(0).all()\nassert ev.target_text.ne(ev.text.map(norm_bn)).all(), (\n    \"A target equals the source text. Preserve the log and inspect the original target procedure.\"\n)\nassert (ev.target_tokens - ev.source_tokens).abs().le(2).all()\nTARGETS = ev[['utt_id', 'text', 'target_text', 'source_tokens', 'target_tokens', 'duration']].copy()\nTARGETS.to_csv(Path(WORK, 'task_b_targets_n100.csv'), index=False)\nPROV['targets_sha256'] = sha256_file(Path(WORK, 'task_b_targets_n100.csv'))\nPROV['tokenizer_sha256'] = hashlib.sha256(json.dumps(\n    processor.tokenizer.get_vocab(), sort_keys=True, ensure_ascii=False\n).encode('utf-8')).hexdigest()\nprint('TARGET PREFLIGHT PASS: 100 targets, original order retained.')\nprint('Targets SHA-256 (must match all six runs):', PROV['targets_sha256'])\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"Validation and explicit resume support embedded in the n100 notebook.\"\"\"\ndef validate_result_rows(frame, expected, rung, checkpoint):\n    required = {'utt_id', 'rung', 'checkpoint', 'edit_type', 'source_text',\n                'target_text', 'source_tokens', 'target_tokens', 'duration',\n                'eps50', 'censored', 'exact_match', 'decoded', 'use_amp'}\n    assert required <= set(frame.columns), 'Result columns missing.'\n    assert len(frame) > 0 and frame.utt_id.notna().all() and frame.utt_id.is_unique\n    assert set(frame.utt_id) <= set(expected.utt_id), 'Unexpected result IDs.'\n    assert frame.rung.eq(rung).all() and frame.checkpoint.eq(checkpoint).all()\n    assert frame.edit_type.eq('T5').all() and frame.use_amp.eq(True).all()\n    assert frame.censored.isin([True, False]).all() and frame.exact_match.isin([True, False]).all()\n    ref = expected.set_index('utt_id').loc[frame.utt_id]\n    for result_col, source_col in [('source_text', 'text'), ('target_text', 'target_text'),\n                                    ('source_tokens', 'source_tokens'), ('target_tokens', 'target_tokens')]:\n        assert frame[result_col].tolist() == ref[source_col].tolist(), f'Mismatched {result_col}.'\n    assert np.allclose(frame.duration, ref.duration, rtol=1e-10, atol=1e-10)\n    ok = frame.loc[~frame.censored]\n    assert ok.eps50.notna().all() and np.isfinite(ok.eps50).all()\n    assert ok.eps50.between(EPS_LO * 0.999, EPS_HI * 1.001).all()\n    assert ok.exact_match.all() and ok.decoded.eq(ok.target_text).all(), (\n        'A reported successful attack failed final exact-target verification. Preserve the log and investigate.'\n    )\n    assert frame.loc[frame.censored, 'eps50'].isna().all()\n    assert frame.loc[frame.censored, 'exact_match'].eq(False).all()\n\n\ndef load_checked_resume(dataset_name, inputs, filename, expected, provenance):\n    if dataset_name is None:\n        return []\n    matches = [Path(p) for p in inputs if Path(p).name == dataset_name]\n    assert len(matches) == 1, f'Attach exactly one resume dataset named {dataset_name}.'\n    csv = one_file(matches[0], filename)\n    previous = json.loads(one_file(matches[0], '_provenance.json').read_text(encoding='utf-8'))\n    for key in ['notebook', 'source_notebook_sha256', 'adapter_sha256', 'attack_eval_sha256',\n                'source_file_sha256', 'targets_sha256', 'tokenizer_sha256', 'model_identity_sha256',\n                'rung', 'checkpoint', 'ckpt_dataset', 'n_eval', 'fixed_eps',\n                'seed', 'use_amp', 'bracket', 'bisect_rounds', 'batch_size',\n                'probe_steps', 'confirm_steps', 'stall_patience', 'check_every', 'lr',\n                'torch', 'transformers', 'numpy', 'pandas']:\n        assert previous.get(key) == provenance[key], f'Resume provenance mismatch: {key}'\n    frame = pd.read_csv(csv, dtype={'utt_id': str})\n    validate_result_rows(frame, expected, RUNG, CHECKPOINT)\n    assert len(frame) < N_EVAL, 'This condition is already complete; do not rerun it accidentally.'\n    assert previous.get('results_sha256') == sha256_file(csv), 'Resume result checksum mismatch.'\n    print(f'Resuming {len(frame)} verified rows from {dataset_name}.')\n    return frame.to_dict('records')\n\n\ndef write_progress(results, out, expected, provenance):\n    frame = pd.DataFrame(results)\n    validate_result_rows(frame, expected, RUNG, CHECKPOINT)\n    frame.to_csv(out, index=False)\n    provenance.update(n_utts=len(frame), n_expected=N_EVAL,\n                      complete=(len(frame) == N_EVAL and set(frame.utt_id) == set(expected.utt_id)),\n                      n_censored=int(frame.censored.sum()),\n                      results_sha256=sha256_file(out))\n    Path(WORK, '_provenance.json').write_text(json.dumps(provenance, indent=2), encoding='utf-8')\n\n\ndef model_identity(model_dir):\n    # Hash the checkpoint files once so a resume cannot use different weights under the same dataset name.\n    root = Path(model_dir)\n    files = sorted(p for p in root.iterdir() if p.is_file() and (\n        p.suffix in {'.json', '.safetensors'} or p.name.startswith('pytorch_model')\n    ))\n    assert any(p.suffix in {'.bin', '.safetensors'} for p in files), 'No checkpoint weight files found.'\n    digest = hashlib.sha256()\n    for path in files:\n        digest.update(path.name.encode('utf-8'))\n        with path.open('rb') as stream:\n            for block in iter(lambda: stream.read(8 * 1024 * 1024), b''):\n                digest.update(block)\n    return digest.hexdigest()\n\nPROV['model_identity_sha256'] = model_identity(MODEL_DIR)\nprint('Checkpoint identity SHA-256:', PROV['model_identity_sha256'])\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5 · Batched differentiable front-end and loss","metadata":{}},{"cell_type":"code","source":"SR = 16000\n\ndef load_batch(rows):\n    \"\"\"Length-bucketed batch. Returns padded waveforms + a validity mask.\"\"\"\n    wavs = []\n    for _, r in rows.iterrows():\n        a, _ = librosa.load(r[\"path\"], sr=SR, mono=True)\n        a = np.asarray(a, np.float32)\n        pk = np.abs(a).max()\n        wavs.append(a / max(pk, 1.0))\n    T = max(len(w) for w in wavs)\n    x = torch.zeros(len(wavs), T, dtype=torch.float32)\n    m = torch.zeros(len(wavs), T, dtype=torch.float32)\n    for i, w in enumerate(wavs):\n        x[i, :len(w)] = torch.from_numpy(w); m[i, :len(w)] = 1.0\n    return x.to(DEVICE), m.to(DEVICE)\n\n\ndef frontend_b(x, mask):\n    \"\"\"Masked zero-mean/unit-variance. Computing stats over the PADDING would\n    corrupt short utterances in a mixed-length batch.\"\"\"\n    if DO_NORMALIZE:\n        n = mask.sum(-1, keepdim=True).clamp(min=1)\n        mu = (x * mask).sum(-1, keepdim=True) / n\n        var = (((x - mu) * mask) ** 2).sum(-1, keepdim=True) / n\n        x = (x - mu) / torch.sqrt(var + 1e-7)\n        x = x * mask\n    return x\n\n\ndef logits_b(x, mask):\n    return model(input_values=frontend_b(x, mask),\n                 attention_mask=mask.long()).logits\n\n\ndef ctc_per_sample(x, mask, tgt_pad, tgt_len):\n    \"\"\"Per-sample CTC loss -> lets us mask out finished samples.\"\"\"\n    lg = logits_b(x, mask)\n    logp = F.log_softmax(lg.float(), -1).transpose(0, 1)          # (L,B,V)\n    in_len = model._get_feat_extract_output_lengths(mask.sum(-1).long()).long()\n    in_len = in_len.clamp(max=logp.shape[0])\n    return F.ctc_loss(logp, tgt_pad, in_len, tgt_len,\n                      blank=BLANK_ID, zero_infinity=True, reduction=\"none\")\n\n\ndef decode_b(x, mask):\n    with torch.no_grad():\n        ids = torch.argmax(logits_b(x, mask), -1)\n    return [norm_bn(s) for s in processor.batch_decode(ids)]\n\n\ndef int16_rt(x):\n    return torch.clamp((x * 32767.0).round(), -32768.0, 32767.0) / 32767.0\n\ndef snr_db_b(clean, adv, mask):\n    n = (adv - clean) * mask\n    s = clean * mask\n    return (10 * torch.log10(s.pow(2).sum(-1) / (n.pow(2).sum(-1) + 1e-12))).cpu().numpy()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6 · Batched PGD with per-sample budgets\n\nEvery utterance in the batch carries its **own** eps, its own success flag and its own\nbest-so-far delta. Finished samples are masked out of the loss so the optimiser stops spending\ncapacity on them, and their delta is frozen at the moment of success.","metadata":{}},{"cell_type":"code","source":"def pgd_batch(x, mask, tgt_pad, tgt_len, targets, eps_vec,\n              steps, delta_init=None, lr=LR):\n    \"\"\"eps_vec: (B,) per-sample L-inf budgets.\n       Returns (success bool array, frozen deltas, steps_used).\"\"\"\n    B = x.shape[0]\n    eps_col = eps_vec.view(-1, 1)\n    delta = (torch.zeros_like(x) if delta_init is None\n             else torch.max(torch.min(delta_init.clone(), eps_col), -eps_col))\n    delta.requires_grad_(True)\n    opt = torch.optim.Adam([delta], lr=lr)\n    scaler = torch.amp.GradScaler(\"cuda\", enabled=USE_AMP)\n\n    active = torch.ones(B, dtype=torch.bool, device=DEVICE)\n    frozen = torch.zeros_like(x)\n    success = np.zeros(B, dtype=bool)\n    best_loss = torch.full((B,), float(\"inf\"), device=DEVICE)\n    stall = torch.zeros(B, dtype=torch.long, device=DEVICE)\n\n    for step in range(steps):\n        if not active.any():\n            break\n        x_adv = torch.clamp(x + delta, -1, 1) * mask\n        with torch.amp.autocast(\"cuda\", enabled=USE_AMP):\n            per = ctc_per_sample(x_adv, mask, tgt_pad, tgt_len)\n        loss = (per * active.float()).sum() / active.float().sum().clamp(min=1)\n\n        opt.zero_grad(); scaler.scale(loss).backward()\n        scaler.unscale_(opt)\n        if delta.grad is not None:\n            delta.grad[~active] = 0          # finished samples stop moving\n        scaler.step(opt); scaler.update()\n        with torch.no_grad():\n            delta.data = torch.max(torch.min(delta.data, eps_col), -eps_col)\n\n        with torch.no_grad():                # stall detection\n            improved = per < best_loss - 1e-3\n            best_loss = torch.minimum(best_loss, per.detach())\n            stall = torch.where(improved, torch.zeros_like(stall), stall + 1)\n            give_up = active & (stall > STALL_PATIENCE)\n            active = active & ~give_up\n\n        if (step + 1) % CHECK_EVERY == 0:\n            hyp = decode_b(int16_rt(torch.clamp(x + delta, -1, 1)) * mask, mask)\n            with torch.no_grad():\n                for i in range(B):\n                    if active[i] and hyp[i] == targets[i]:\n                        frozen[i] = delta[i].detach()\n                        success[i] = True\n                        active[i] = False\n    return success, frozen, step + 1","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7 · Per-sample geometric bisection, run in lockstep","metadata":{}},{"cell_type":"code","source":"def eps50_batch(x, mask, tgt_pad, tgt_len, targets, verbose=True):\n    \"\"\"Per-sample geometric bisection for the smallest budget that still works.\n\n    ---------------------------------------------------------------------------\n    BUG FIXED 2026-09-22 (found in the rung-1 T5 run, 40% censoring).\n\n    The previous version ended with\n\n        ok, best, _ = pgd_batch(..., hi, CONFIRM_STEPS, delta_init=best)\n\n    which REASSIGNED `best` wholesale to pgd_batch's return value. That return\n    is a `frozen` buffer holding zeros for any sample that did not succeed on\n    that particular pass -- so one unlucky final pass wiped out a delta that had\n    already been verified at a larger budget. Those samples were then written\n    out as censored, with SNR 70-86 dB (the signature of delta == 0) even though\n    the ceiling probe had already proved them attackable.\n\n    In the rung-1 run that discarded ~6 of 25 real measurements. Worse, it does\n    so at a RATE THAT VARIES BY MODEL, so comparing median eps50 across rungs\n    would have been comparing differently-biased subsamples -- straight into\n    RQ1's headline.\n\n    Fix: keep a running `best_eps` / `best_delta` that is only ever updated on a\n    VERIFIED success. Nothing overwrites a confirmed result. Censoring now means\n    exactly one thing: the attack failed at every budget tried, ceiling included.\n\n    Side benefit: eps50 becomes a guaranteed UPPER BOUND on the true minimum.\n    Optimiser noise can only make it conservative, never wrong in an\n    unpredictable direction -- which is the property you want in a metric you\n    are about to compare across four models.\n    ---------------------------------------------------------------------------\n    \"\"\"\n    B = x.shape[0]\n    lo = torch.full((B,), EPS_LO, device=DEVICE)\n    hi = torch.full((B,), EPS_HI, device=DEVICE)\n    best_eps   = torch.full((B,), float(\"nan\"), device=DEVICE)   # verified only\n    best_delta = torch.zeros_like(x)\n\n    def record(okt, eps_vec, d):\n        \"\"\"Adopt a result only when it is both successful and better.\"\"\"\n        nonlocal best_eps, best_delta\n        better = okt & (torch.isnan(best_eps) | (eps_vec < best_eps))\n        best_eps   = torch.where(better, eps_vec, best_eps)\n        best_delta = torch.where(better.view(-1, 1), d, best_delta)\n\n    # Ceiling probe. Failure here is genuine right-censoring.\n    ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, hi, CONFIRM_STEPS)\n    okt = torch.tensor(ok, device=DEVICE)\n    record(okt, hi, d)\n    feasible = ok.copy()\n    if verbose:\n        print(f\"  ceiling probe eps={EPS_HI:.1e}: {feasible.sum()}/{B} feasible, \"\n              f\"{(~feasible).sum()} censored\")\n\n    for rnd in range(BISECT_ROUNDS):\n        mid = torch.sqrt(lo * hi)\n        ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, mid,\n                             PROBE_STEPS, delta_init=best_delta)\n        ok = ok & feasible\n        okt = torch.tensor(ok, device=DEVICE)\n        hi = torch.where(okt, mid, hi)\n        lo = torch.where(okt, lo, mid)\n        record(okt, mid, d)\n        if verbose:\n            print(f\"  round {rnd+1}/{BISECT_ROUNDS}: {ok.sum()}/{B} succeeded \"\n                  f\"| median eps {mid.median().item():.3e}\")\n\n    # Budget-consistency round: one more bisection at the FULL step budget, so\n    # the reported boundary is measured at the budget eps50 is defined at rather\n    # than at the cheaper probe budget. See RUNLOG for why this matters.\n    #\n    # Skipped in FIXED_EPS mode: there is no bracket to refine. The ceiling probe\n    # already ran at CONFIRM_STEPS, so the single reported budget is measured at\n    # the full budget and needs no second pass.\n    if BISECT_ROUNDS == 0:\n        eps       = best_eps.cpu().numpy()\n        confirmed = ~np.isnan(eps)\n        if verbose:\n            print(f\"  -> FIXED_EPS {EPS_HI:.1e}: {confirmed.sum()}/{B} succeeded, \"\n                  f\"{(~confirmed).sum()} resisted\")\n        return eps, np.ones(B), confirmed, best_delta\n\n    mid = torch.sqrt(lo * hi)\n    ok, d, _ = pgd_batch(x, mask, tgt_pad, tgt_len, targets, mid,\n                         CONFIRM_STEPS, delta_init=best_delta)\n    ok = ok & feasible\n    okt = torch.tensor(ok, device=DEVICE)\n    record(okt, mid, d)\n    if verbose:\n        print(f\"  budget-consistency round @ {CONFIRM_STEPS} steps: {ok.sum()}/{B}\")\n\n    eps       = best_eps.cpu().numpy()\n    confirmed = ~np.isnan(eps)                 # censored == never succeeded\n    width     = (hi / lo).cpu().numpy()\n    if verbose:\n        print(f\"  -> {confirmed.sum()}/{B} measured, {(~confirmed).sum()} censored\")\n    return eps, width, confirmed, best_delta\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8 · Validate the batching before trusting it\n\nBatching is a substantial rewrite. If masked normalisation or the per-sample CTC lengths are\nwrong, results would be quietly biased rather than obviously broken. So: run two utterances\nbatched, then the same two alone, and compare.\n\n> **This costs about three extra eps50 searches (~1–2 GPU-h).** Run it once, for\n> `(RUNG=1, EDIT_TYPE=\"T5\")`, record the agreement in `RUNLOG.md`, then set\n> `VALIDATE_BATCHING = False` for the other seven runs. It is a one-off correctness check on\n> the batching code, and that code does not change between runs.","metadata":{}},{"cell_type":"code","source":"if VALIDATE_BATCHING and len(ev) >= 2:\n    sub = ev.head(2)\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    print(\"batched (2 together):\")\n    t0 = time.time(); e_b, _, c_b, _ = eps50_batch(xb, mb, tpad, tl, tgts); tb = time.time()-t0\n\n    print(\"\\nsingle (one at a time):\")\n    e_s = []\n    t0 = time.time()\n    for i in range(2):\n        xs, ms = xb[i:i+1], mb[i:i+1]\n        e, _, c, _ = eps50_batch(xs, ms, tpad[i:i+1], tl[i:i+1], [tgts[i]])\n        e_s.append(e[0])\n    ts = time.time()-t0\n\n    print(f\"\\n{'':6s} {'batched':>12s} {'single':>12s}  {'rel diff':>9s}\")\n    for i in range(2):\n        a, b = e_b[i], e_s[i]\n        rd = \"n/a\" if (np.isnan(a) or np.isnan(b)) else f\"{100*abs(a-b)/b:8.1f}%\"\n        print(f\"  #{i}  {a:12.3e} {b:12.3e}  {rd:>9s}\")\n    print(f\"\\nwall-clock: batched {tb:.0f}s vs single {ts:.0f}s -> {ts/tb:.2f}x\")\n    print(\"Agreement within ~15% is fine (PGD is stochastic). Larger means a batching bug.\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8b · Is eps50 converged at this step budget?\n","metadata":{}},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\n# CONVERGENCE CHECK.\n#\n# eps50 is defined relative to CONFIRM_STEPS. If doubling the budget moves it\n# materially, the metric is still tracking optimiser convergence rather than\n# model robustness -- and since rungs may converge at different rates, that\n# would leak straight into the RQ1 comparison. Three utterances is enough to\n# detect it.\n# ---------------------------------------------------------------------------\nif len(ev) >= 3:\n    sub = ev.head(3)\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    _saved = CONFIRM_STEPS\n    print(f\"at CONFIRM_STEPS = {_saved}:\")\n    e1, _, _, _ = eps50_batch(xb, mb, tpad, tl, tgts, verbose=False)\n    CONFIRM_STEPS = _saved * 2\n    print(f\"at CONFIRM_STEPS = {CONFIRM_STEPS}:\")\n    e2, _, _, _ = eps50_batch(xb, mb, tpad, tl, tgts, verbose=False)\n    CONFIRM_STEPS = _saved\n\n    print(f\"\\n{'utt':>4s} {'@'+str(_saved):>12s} {'@'+str(_saved*2):>12s} {'shift':>9s}\")\n    shifts = []\n    for i in range(len(sub)):\n        a, b = e1[i], e2[i]\n        if np.isnan(a) or np.isnan(b):\n            print(f\"{i:>4d} {'censored':>12s}\"); continue\n        sh = abs(a - b) / a; shifts.append(sh)\n        print(f\"{i:>4d} {a:12.3e} {b:12.3e} {100*sh:8.1f}%\")\n    if shifts:\n        ms = float(np.median(shifts))\n        print(f\"\\nmedian shift on doubling the budget: {100*ms:.1f}%\")\n        if ms < 0.10:\n            print(\"Converged. eps50 is stable at this budget -- safe to compare rungs.\")\n        else:\n            print(\"NOT CONVERGED. Raise CONFIRM_STEPS and re-run, or state explicitly\")\n            print(\"that eps50 is budget-relative and hold the budget fixed across ALL\")\n            print(\"rungs so the comparison stays internally valid.\")\n        PROV[\"convergence_shift_pct\"] = round(100 * ms, 1)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9 · The main sweep","metadata":{}},{"cell_type":"code","source":"ev = ev.sort_values(\"duration\").reset_index(drop=True)   # length buckets = less padding\nOUT = f\"{WORK}/eps50_rung{RUNG}_{CHECKPOINT}_{EDIT_TYPE}_n100.csv\"\n\nassert_task_b_settings()\nresults = load_checked_resume(RESUME_DATASET, inputs, Path(OUT).name, ev, PROV)\nif results:\n    assert [r['utt_id'] for r in results] == ev.utt_id.tolist()[:len(results)], (\n        'Resume rows must be the original sorted batch prefix.'\n    )\ndone = {r[\"utt_id\"] for r in results}\ntodo = ev[~ev.utt_id.isin(done)].reset_index(drop=True)\n\nbatches = [todo.iloc[i:i+BATCH_SIZE] for i in range(0, len(todo), BATCH_SIZE)]\nprint(f\"{len(todo)} utterances remaining -> {len(batches)} batches of up to {BATCH_SIZE}\\n\")\n\nDEADLINE = time.time() + SESSION_HOURS * 3600\nt_start = time.time()\n\nfor bi, sub in enumerate(batches):\n    if time.time() > DEADLINE:\n        print(f\"*** SESSION BUDGET ({SESSION_HOURS} h) reached. \"\n              f\"{sum(len(b) for b in batches[bi:])} utterances not yet measured. ***\\n\")\n        break\n\n    print(f\"=== batch {bi+1}/{len(batches)} ({len(sub)} utts, \"\n          f\"{sub.duration.min():.1f}-{sub.duration.max():.1f}s) ===\")\n    xb, mb = load_batch(sub)\n    tp = [torch.tensor(processor.tokenizer(t).input_ids, device=DEVICE) for t in sub.target_text]\n    tl = torch.tensor([len(t) for t in tp], device=DEVICE)\n    tpad = torch.nn.utils.rnn.pad_sequence(tp, batch_first=True, padding_value=0)\n    tgts = list(sub.target_text)\n\n    t0 = time.time()\n    eps, width, conf, best = eps50_batch(xb, mb, tpad, tl, tgts)\n    dt = time.time() - t0\n\n    x_adv = int16_rt(torch.clamp(xb + best, -1, 1)) * mb\n    snrs = snr_db_b(xb, x_adv, mb)\n    hyp = decode_b(x_adv, mb)\n\n    for i, (_, r) in enumerate(sub.iterrows()):\n        results.append(dict(\n            utt_id=r[\"utt_id\"], rung=RUNG, edit_type=EDIT_TYPE, checkpoint=CHECKPOINT,\n            duration=r[\"duration\"], source_text=r[\"text\"], target_text=r[\"target_text\"],\n            source_tokens=int(r[\"source_tokens\"]), target_tokens=int(r[\"target_tokens\"]),\n            eps50=float(eps[i]) if not np.isnan(eps[i]) else None,\n            # SNR/decode are meaningful only where an attack was verified.\n            # Writing them for censored rows produced the 70-86 dB values in\n            # the rung-1 run, which looked like data but were delta == 0.\n            censored=bool(np.isnan(eps[i])),\n            eps50_per_token=(float(eps[i]) / r[\"target_tokens\"]) if not np.isnan(eps[i]) else None,\n            bracket_width=float(width[i]), snr_db=float(snrs[i]),\n            decoded=(hyp[i] if not np.isnan(eps[i]) else None),\n            exact_match=(bool(hyp[i] == r[\"target_text\"])\n                         if not np.isnan(eps[i]) else False),\n            use_amp=USE_AMP, model_dir=MODEL_DIR,\n        ))\n\n    # Write after EVERY batch. The original only wrote at the very end, so a\n    # session killed on the last batch threw away hours of finished work.\n    write_progress(results, OUT, ev, PROV)\n\n    remaining = len(batches) - (bi + 1)\n    print(f\"  batch done in {dt/60:.1f} min | {(~np.isnan(eps)).sum()}/{len(sub)} uncensored \"\n          f\"| {len(results)} rows written\")\n    if remaining:\n        print(f\"  projected for the {remaining} remaining batches: \"\n              f\"{remaining*dt/3600:.1f} h \"\n              f\"({'fits' if time.time()+remaining*dt < DEADLINE else 'WILL NOT FIT this session'})\\n\")\n    else:\n        print()\n\nprint(f\"TOTAL {(time.time()-t_start)/60:.1f} min this session | {len(results)}/{len(ev)} measured\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 10 · Results and save","metadata":{}},{"cell_type":"code","source":"CK_TAG = \"final\" if CHECKPOINT == \"final\" else \"ms\" + CHECKPOINT.split(\"-\")[1]\nif FIXED_EPS is not None:\n    CK_TAG += \"-fixedeps\"        # keep trajectory runs separate from eps50 runs\n# Every checkpoint of a rung gets its OWN dataset. Before this, 100h final and\n# all three 100h milestones were told to publish as \"eps50-rung100-T5\", so\n# \"New Version\" on that dataset silently replaced one checkpoint's results\n# with another's.\ndf = pd.DataFrame(results)\nok = df[~df.censored] if len(df) else df\nvalidate_result_rows(df, ev, RUNG, CHECKPOINT)\ncomplete = len(df) == N_EVAL and set(df.utt_id) == set(EXPECTED_IDS)\n\n# Clean WER on the SAME attack utterances, measured here so a trajectory run\n# needs no NB-02 pass. The figure plots attack success against clean WER and both\n# numbers must come from the same checkpoint.\nif FIXED_EPS is not None:\n    try:\n        import jiwer\n        _wav, _msk = load_batch(ev)\n        _hyp = decode_b(_wav, _msk)\n        _ref = [norm_bn(t) for t in ev.text]\n        _pr = [(r, h) for r, h in zip(_ref, _hyp) if r.strip()]\n        PROV[\"clean_wer_attackset\"] = float(jiwer.wer([p[0] for p in _pr], [p[1] for p in _pr]))\n        PROV[\"clean_cer_attackset\"] = float(jiwer.cer([p[0] for p in _pr], [p[1] for p in _pr]))\n        print(f\"clean WER on the attack set: {PROV['clean_wer_attackset']:.4f}  \"\n              f\"CER {PROV['clean_cer_attackset']:.4f}   <- x-axis for the trajectory figure\")\n        del _wav, _msk; torch.cuda.empty_cache()\n    except Exception as e:\n        print(f\"clean-WER measurement failed ({type(e).__name__}: {e}); \"\n              f\"fall back to NB-02's clean_wer_all_checkpoints.csv\")\n\nprint(f\"=== RUNG {RUNG}h | {EDIT_TYPE}\"\n      + (f\" | FIXED_EPS {FIXED_EPS:.1e} ===\" if FIXED_EPS is not None else \" ===\"))\nif FIXED_EPS is not None:\n    print(\"  FIXED_EPS mode: every success is recorded at exactly the fixed budget, so\")\n    print(\"  eps50, its IQR and its bootstrap CI below are DEGENERATE and carry no\")\n    print(\"  information. Use ASR@eps and the per-utterance `censored` column only.\")\nprint(f\"utterances      : {len(df)} of {len(ev)}\")\nprint(f\"censored        : {df.censored.sum()} ({100*df.censored.mean():.0f}%)  <- report this\")\nif len(ok):\n    print(f\"median eps50    : {ok.eps50.median():.3e}  ({ok.eps50.median()*32767:.1f} int16 units)\")\n    print(f\"IQR             : {ok.eps50.quantile(.25):.3e} - {ok.eps50.quantile(.75):.3e}\")\n    print(f\"per target token: {ok.eps50_per_token.median():.3e}   <- length-normalised\")\n    print(f\"median SNR      : {ok.snr_db.median():.1f} dB\")\n    print(f\"exact match     : {ok.exact_match.sum()}/{len(ok)}\")\n\n    # A bracket width near 1.0 with eps50 at the floor means the search hit\n    # EPS_LO: LEFT-censored, and just as biasing as the right-censored kind.\n    at_floor = int(((ok.eps50 <= EPS_LO * 1.05)).sum())\n    if at_floor:\n        print(f\"WARNING: {at_floor} utterances sit at the EPS_LO floor ({EPS_LO:.0e}) \"\n              \"-- left-censored. Record this warning; retain the shared Task B budget.\")\n\n    b = [ok.eps50.sample(len(ok), replace=True).median() for _ in range(2000)]\n    print(f\"bootstrap 95% CI: [{np.percentile(b,2.5):.3e}, {np.percentile(b,97.5):.3e}]\")\n\ndf.to_csv(OUT, index=False)\nPROV.update(n_utts=len(df), n_expected=len(ev), complete=bool(complete),\n            n_censored=int(df.censored.sum()) if len(df) else 0,\n            median_eps50=float(ok.eps50.median()) if len(ok) else None,\n            median_eps50_per_token=float(ok.eps50_per_token.median()) if len(ok) else None,\n            median_snr=float(ok.snr_db.median()) if len(ok) else None,\n            minutes=round((time.time()-t_start)/60, 1))\njson.dump(PROV, open(f\"{WORK}/_provenance.json\", \"w\"), indent=2)\n\nKAGGLE_USER = os.environ.get(\"KAGGLE_USERNAME\") or \"YOURUSERNAME\"\njson.dump({\"title\": f\"eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}-n100\",\n           \"id\": f\"{KAGGLE_USER}/eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}-n100\",\n           \"licenses\": [{\"name\": \"CC0-1.0\"}]},\n          open(f\"{WORK}/dataset-metadata.json\", \"w\"), indent=2)\n\nif not complete:\n    print(\"\\n\" + \"=\" * 70)\n    print(f\"  PARTIAL: {len(df)}/{len(ev)} utterances measured. Nothing is lost.\")\n    print(f\"    1. Output panel -> Create Dataset (or New Version) \"\n          f\"-> eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}-n100\")\n    print( \"    2. Attach that dataset back as an INPUT to this notebook\")\n    print( \"    3. Set RESUME_DATASET to that exact name, keep RUN_ID, and re-commit in a fresh session.\")\n    print(\"=\" * 70)\nelse:\n    print(f\"\\nRUNG {RUNG}h / {EDIT_TYPE} COMPLETE -- publish as \"\n          f\"eps50-rung{RUNG}-{CK_TAG}-{EDIT_TYPE.lower()}-n100 and record the median in RUNLOG.md.\")\n\nprint(\"\\n\" + json.dumps(PROV, indent=2))\ndf.head(10)\n\n# ---------------------------------------------------------------------------\n# CO-PRIMARY METRIC: attack success rate at a FIXED budget.\n#\n# Median eps50 is computed only over UNCENSORED utterances, so if the censoring\n# rate differs between rungs -- and it will -- the medians describe different\n# subpopulations and are not directly comparable. That is a censoring bias\n# sitting directly under RQ1's headline.\n#\n# ASR@ceiling has no such problem: every utterance contributes, success or\n# failure, so it is comparable across rungs by construction.\n#\n# The two measure different things and should both be reported:\n#   ASR@ceiling -> can this model be STEERED to an exact target at all?\n#   median eps50 -> when it can, how much perturbation does it take?\n#\n# Expect them to move in opposite directions. A better model is usually easier\n# to aim (lower censoring) but harder to push (higher eps50). Reporting only\n# one of them would tell half the story, and the flattering half depends on\n# which one you pick.\n# ---------------------------------------------------------------------------\nASR_AT_CEILING = float((~df.censored).mean())\nprint(f\"\\nASR@eps={EPS_HI:.0e}   : {ASR_AT_CEILING:.1%}  ({(~df.censored).sum()}/{len(df)})\")\nprint(f\"  <- censoring-free, uses all {len(df)} utterances, comparable across rungs\")\nprint(f\"median eps50 is over the {(~df.censored).sum()} steerable utterances only\")\n\n# ---------------------------------------------------------------------------\n# LEFT-censoring: the mirror image of the block above, and the one this notebook\n# had no check for. If eps50 lands at the bottom of the bracket, the attack\n# succeeded at the smallest budget searched, so the value is an UPPER BOUND and\n# the true eps50 is somewhere below it. Reporting it as a point estimate\n# understates how cheap the attack is -- in the direction that flatters RQ2.\n# ---------------------------------------------------------------------------\nfloor_hits = int((ok.eps50 <= EPS_LO * 1.15).sum()) if len(ok) else 0\nPROV[\"left_censored\"] = floor_hits\nif floor_hits:\n    print(f\"\\n*** LEFT-CENSORED: {floor_hits}/{len(ok)} eps50 values sit at the \"\n          f\"bracket floor ({EPS_LO:.1e}) ***\")\n    print(\"    These are upper bounds. Record as censored-below; retain Task B settings.\")\n\nif len(ok):\n    b = [ok.eps50.sample(len(ok), replace=True).median() for _ in range(2000)]\n    lo_ci, hi_ci = np.percentile(b, 2.5), np.percentile(b, 97.5)\n    print(f\"\\neps50 bootstrap 95% CI spans {hi_ci/lo_ci:.2f}x\")\n    if hi_ci / lo_ci > 1.8:\n        print(\"  WIDE. Between-rung differences must exceed this to be credible.\")\n        print(\"  Record the wide interval; retain the shared Task B budget.\")\n\nPROV[\"asr_at_ceiling\"] = round(ASR_AT_CEILING, 4)\nPROV[\"eps50_ci_ratio\"] = round(float(hi_ci / lo_ci), 3) if len(ok) else None\n\n# Save AFTER all statistics are added. Earlier NB-05 wrote provenance too soon.\nPROV['censoring_rate'] = float(df.censored.mean())\nPROV['median_eps50_successes_only'] = float(ok.eps50.median()) if len(ok) else None\nPROV['minutes_including_setup_and_diagnostics'] = round((time.time() - NOTEBOOK_STARTED) / 60, 1)\nwrite_progress(results, OUT, ev, PROV)\nprint('FINAL AUDIT:', len(df), '/', N_EVAL, 'unique rows; complete =', PROV['complete'])\nprint('Publish as:', OUTPUT_DATASET)\nprint('Report censoring, convergence, and floor warnings; keep all rows.')\nprint('Record notebook/input/output versions and targets SHA-256 in RUNLOG.')\nprint('The lead author handles the established censoring-aware aggregation and pair effects.')\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Before the next condition\n\n- Require 100 unique expected IDs and `complete: true`; partial outputs are not finished conditions.\n- Publish the CSV, `_provenance.json`, and `task_b_targets_n100.csv` as this condition's **new n100 dataset**.\n- Preserve the notebook log; record censoring, convergence, floor warnings, target hash and dataset versions in RUNLOG.\n- Compare the target hash against the first run. All six conditions must attack the same targets.\n- Change RUN_ID only for the next fresh condition. Clear RESUME_DATASET to None.\n- Do not update historical n25 datasets. Do not report the successes-only median as a censoring-aware median.\n- Six datasets plus six experiment RUNLOG entries go back to the lead author for aggregation.\n","metadata":{}}]}