{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"3023a8c4-3464-4d7f-8e85-aad3a39d4fc5","cell_type":"markdown","source":"# NB-11 — Whisper Arm (Task C)\n\n**One question:** does a *different architecture* — autoregressive encoder–decoder, not CTC —\ncost more or less to attack than our XLS-R ladder?\n\n**What this is NOT.** Whisper's training data is not ours and is not controlled. This is **one\nhonest datapoint on architecture dependence**, not a controlled comparison. The paper must say\nso in the same sentence that reports the number.\n\n---\n\n## What you will get\n\n| | |\n|---|---|\n| Whisper clean WER on the attack set | gate — if it cannot transcribe Bengali, stop |\n| **ASR at eps = 2e-3** | the ladder's own ceiling, directly comparable |\n| ASR at larger eps | how much budget Whisper actually needs |\n| **Transfer**: do the ladder's adversarial WAVs fool Whisper? | free, the WAVs already exist |\n\n## Comparable to\n\n`ASR@2e-3` from Task B, same 25 or 100 utterances, **byte-identical targets** (read from the\nladder's own `task_b_targets_n100.csv`, never regenerated):\n\n| condition | ASR@2e-3, n=100 |\n|---|---|\n| 1 h final | 75% |\n| 5 h final | 79% |\n| 25 h final | 66% |\n| 100 h ms-1000 | 37% |\n| 100 h ms-2000 | 46% |\n| 100 h ms-3500 | 53% |\n\n## The two traps, both now guarded in code\n\n1. **The front end.** HuggingFace's Whisper feature extractor goes through numpy, so the gradient\n   never reaches the waveform. Cell 5 reimplements the log-mel in torch and **asserts** it matches\n   HF to 1e-4 and that `waveform.grad` is finite and non-zero. Verified locally against\n   `transformers==4.44.2`: max abs difference **1.19e-07**.\n2. **Language forcing.** Without a pinned decoder prompt Whisper silently translates to English.\n   Cell 4 pins `<|bn|><|transcribe|><|notimestamps|>` and **prints clean transcriptions for you to\n   look at** before anything long runs.\n\n## Setup on Kaggle\n\n* **Internet must be ON** (Settings → Internet). Whisper downloads from the HF hub.\n* Accelerator **GPU T4 x2**.\n* Attach: `bengaliai-speech` (audio), `bn-asr-splits` (eval CSV), **one** `eps50-*-n100` dataset\n  (for the frozen targets), and for the transfer stage `pipeline-survival-rung100`.\n\n## If it fails\n\nSay where it stopped. *\"We attempted a Whisper arm; the differentiable front end was completed but\nthe attack did not reach the ladder's budget\"* is a reportable sentence. A rushed number is worse\nthan no number.","metadata":{}},{"id":"3e5f34ae-0399-49fe-bd52-0128b75cc5ee","cell_type":"markdown","source":"## 1 · Configuration — one knob","metadata":{}},{"id":"49d1a30e-2732-4fe3-bd59-60345e6eed92","cell_type":"code","source":"STAGE = \"calibrate\"     # <<<< \"calibrate\" -> \"attack\" -> \"transfer\"\n                        #      run them in that order; calibrate is ~10 GPU-min\n                        #      and tells you whether \"attack\" is worth 1-3 GPU-h.\n\n# ---- everything below is locked; changing it breaks comparability --------------\nMODEL_ID    = \"openai/whisper-large-v3\"   # measured 2026-09-29 by NB-11b: this is the\n                        # ONLY Whisper that emits Bengali at all. small gives 0% Bengali\n                        # (64% Devanagari), medium hallucinates repetition loops.\nN_EVAL      = 25         # 25 = the RQ1 spine. 100 = the full Task B set (~4x runtime).\nEDIT_TYPE   = \"T5\"       # the ladder's T5 targets; T1 is not defined for this arm\nEPS_MAIN    = 2.0e-3     # the ladder's ceiling -- the headline number\nEPS_EXTRA   = [8.0e-3, 3.2e-2]   # how much budget Whisper needs if 2e-3 fails\nSEED        = 1337\nSR          = 16000\nBATCH_SIZE  = 1          # large-v3 is 1.55B params and every clip is padded to 30 s.\n                         # Raise only if calibration reports headroom.\nGRAD_CKPT   = True       # needed to fit large-v3 backward on a T4\nATTACK_STEPS= 400\nLR          = 1.0e-4     # calibration may tell you to change this -- it is the one\n                         # number you are allowed to revise, and only from evidence\nCHECK_EVERY = 100        # generation with large-v3 is slow; checking less often costs\n                         # resolution on hit_step, not attack strength\nUSE_AMP     = True\nMAX_NEW     = None       # DERIVED in cell 5 from the real targets. Do not hardcode:\n                         # measured, the longest of the 100 ladder targets is 202\n                         # Whisper tokens (the ladder's own char vocab calls it 101),\n                         # so any guessed cap truncates generation and silently\n                         # converts a success into a failure.\n\n# MODEL_ID is deliberately NOT locked -- which Whisper can read Bengali is an\n# empirical question NB-11b answered. Everything that governs comparability is.\n_LOCKED = dict(edit=EDIT_TYPE, eps=EPS_MAIN, seed=SEED, sr=SR)\nassert _LOCKED == dict(edit=\"T5\", eps=2.0e-3, seed=1337, sr=16000), (\n    \"Locked settings were edited. The ASR@2e-3 comparison against the ladder is only \"\n    \"valid at eps=2e-3 with T5 targets. If you meant to explore, use EPS_EXTRA.\")\nassert STAGE in (\"calibrate\", \"attack\", \"transfer\"), f\"unknown STAGE {STAGE!r}\"\nassert N_EVAL in (25, 100), \"N_EVAL must be 25 (RQ1 spine) or 100 (full Task B set)\"\n\nprint(f\"NB-11 Whisper arm | STAGE = {STAGE} | {MODEL_ID} | n = {N_EVAL}\")\nprint(f\"  eps_main {EPS_MAIN:.1e}   steps {ATTACK_STEPS}   lr {LR:.1e}   batch {BATCH_SIZE}\")\nif STAGE == \"calibrate\":\n    print(\"  -> calibration only. No headline number is produced by this stage.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"bf93912d-802e-4dce-98ac-0e7479e37caf","cell_type":"markdown","source":"## 2 · Boilerplate","metadata":{}},{"id":"5e2d1469-77ea-488c-92ac-a9f9190aa97a","cell_type":"code","source":"import os, re, json, math, time, unicodedata, hashlib, glob\nimport numpy as np, pandas as pd, torch, torch.nn.functional as F\nimport transformers          # here, NOT in the librosa cell: the provenance block\n                             # at the end reads transformers.__version__, and if the\n                             # librosa import ever fails that name would be undefined.\n\nZW = dict.fromkeys(map(ord, \"\\u200c\\u200d\"), None)\ndef norm_bn(s):\n    \"\"\"The project-wide Bengali normaliser. Defined here because cell 5 needs it\n    before the loss cell runs.\"\"\"\n    return re.sub(r\"\\s+\", \" \", unicodedata.normalize(\"NFC\", str(s)).translate(ZW)).strip()\n\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\ntorch.manual_seed(SEED); np.random.seed(SEED)\nprint(\"torch\", torch.__version__, \"|\", DEVICE,\n      torch.cuda.get_device_name(0) if DEVICE == \"cuda\" else \"\")\nif DEVICE == \"cpu\":\n    print(\"  WARNING: no GPU. The attack stage will take days on CPU. Calibrate only.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"3455f482-f9d7-4cf4-90e4-80389b2b2727","cell_type":"code","source":"# transformers 4.44.2 is what the whole project pins. Whisper's generate() API\n# changed around 4.45 (forced_decoder_ids deprecation), so do not float this.\nprint(\"transformers\", transformers.__version__)\nif transformers.__version__ != \"4.44.2\":\n    print(\"  NOTE: project pins 4.44.2. If generate() errors on language=/task=, \"\n          \"run:  !pip install -q transformers==4.44.2   and restart the kernel.\")\nimport librosa\nprint(\"librosa\", librosa.__version__)","metadata":{},"outputs":[],"execution_count":null},{"id":"66b0913e-9c43-459a-85e2-5e65da3ca812","cell_type":"markdown","source":"## 3 · Resolve inputs, eval set and the **frozen** targets","metadata":{}},{"id":"c24a06e2-085a-421b-9a9a-cbe704ddcabf","cell_type":"code","source":"# Your own datasets mount THREE levels deep at /kaggle/input/datasets/<user>/<name>,\n# so a flat glob of /kaggle/input/* sees only wrapper dirs. That mistake silently\n# switched the eval set once already and cost four runs. Resolve explicitly.\ninputs = []\nfor pat in (\"/kaggle/input/*\", \"/kaggle/input/*/*\", \"/kaggle/input/*/*/*\"):\n    inputs += [p for p in glob.glob(pat) if os.path.isdir(p)]\ninputs = sorted(set(inputs))\nprint(\"visible input dirs:\", len(inputs))\n\ndef find_file(name):\n    for p in inputs:\n        c = os.path.join(p, name)\n        if os.path.exists(c):\n            return c\n    return None\n\nSPLITS = next((p for p in inputs if os.path.exists(f\"{p}/attack_eval_25.csv\")), None)\nassert SPLITS, (\"bn-asr-splits not attached (looking for attack_eval_25.csv). \"\n                f\"Visible: {[os.path.basename(p) for p in inputs][:20]}\")\n\nTGT_CSV = find_file(\"task_b_targets_n100.csv\")\nassert TGT_CSV, (\n    \"task_b_targets_n100.csv not found.\\\\n\"\n    \"Attach ONE of the Task B result datasets, e.g.\\\\n\"\n    \"  zoayriaabedin/eps50-rung100-milestone-3500-t5-n100\\\\n\"\n    \"The targets MUST come from there. Regenerating them with Whisper's tokenizer \"\n    \"would give different sentences and the comparison against the ladder would be void.\")\nprint(\"targets  :\", TGT_CSV)\n\ntg = pd.read_csv(TGT_CSV)\nassert len(tg) == 100 and tg.utt_id.nunique() == 100, f\"target file has {len(tg)} rows\"\nTGT_SHA = hashlib.sha256(open(TGT_CSV, \"rb\").read()).hexdigest()\nprint(\"targets sha256:\", TGT_SHA)\nif TGT_SHA != \"8711dc56202b33f90e769c9735231b0fa8cf27fcbc20151f5466bcd9cc2b1256\":\n    print(\"  !! this is NOT the hash the six ladder runs used. Comparability is broken.\")\nelse:\n    print(\"  matches all six ladder runs -- byte-identical targets confirmed\")\n\n# expected_ids order: original 25 first. Row order in the RESULT csvs is duration-sorted,\n# but this target file is in manifest order, so head(25) IS the RQ1 spine.\nev = tg.head(N_EVAL).copy().reset_index(drop=True)\nprint(f\"eval set: {len(ev)} utterances \"\n      f\"({'RQ1 spine' if N_EVAL == 25 else 'full Task B set'})\")","metadata":{},"outputs":[],"execution_count":null},{"id":"9f7e1732-11e5-4d63-bbc5-9ebac521c886","cell_type":"code","source":"# Audio. bengaliai-speech is COMPETITION data: you must join the competition and\n# accept its rules before it will even appear under \"+ Add Input\".\nAUDIO_ROOT = None\nfor cand in (\"/kaggle/input/competitions/bengaliai-speech\",\n             \"/kaggle/input/bengaliai-speech\"):\n    if os.path.isdir(cand):\n        AUDIO_ROOT = cand; break\nassert AUDIO_ROOT, (\"bengaliai-speech not attached. It is competition data -- join the \"\n                    \"competition and accept the rules first, then + Add Input.\")\n\ndef audio_path(uid):\n    for sub in (\"train_mp3s\", \"examples\"):\n        p = f\"{AUDIO_ROOT}/{sub}/{uid}.mp3\"\n        if os.path.exists(p): return p\n    return None\n\nev[\"path\"] = ev.utt_id.map(audio_path)\nmissing = ev[ev.path.isna()]\nassert len(missing) == 0, f\"{len(missing)} audio files not found, e.g. {missing.utt_id.tolist()[:5]}\"\nprint(f\"all {len(ev)} audio files present under {AUDIO_ROOT}\")","metadata":{},"outputs":[],"execution_count":null},{"id":"91a1f0e7-3f48-4642-9fe2-2690b9c45896","cell_type":"markdown","source":"## 4 · Load Whisper and **pin the language** (trap 2)","metadata":{}},{"id":"6e20698c-9fed-431b-b0c6-89ff9b844145","cell_type":"code","source":"from transformers import (WhisperForConditionalGeneration, WhisperProcessor,\n                          WhisperFeatureExtractor)\n\nmodel = WhisperForConditionalGeneration.from_pretrained(MODEL_ID).to(DEVICE).eval()\nproc  = WhisperProcessor.from_pretrained(MODEL_ID)\nfe    = WhisperFeatureExtractor.from_pretrained(MODEL_ID)\ntok   = proc.tokenizer\nfor p in model.parameters():\n    p.requires_grad_(False)          # we optimise the waveform, never the weights\n\n# generation_config ships forced_decoder_ids that CONFLICT with language=/task= and\n# make generate() print a warning and ignore one of them. Clear both, then force\n# explicitly at every call site so there is exactly one source of truth.\nmodel.generation_config.forced_decoder_ids = None\nmodel.config.forced_decoder_ids = None\n\nSOT = tok.convert_tokens_to_ids(\"<|startoftranscript|>\")\nBN  = tok.convert_tokens_to_ids(\"<|bn|>\")\nTR  = tok.convert_tokens_to_ids(\"<|transcribe|>\")\nNT  = tok.convert_tokens_to_ids(\"<|notimestamps|>\")\nEOT = tok.convert_tokens_to_ids(\"<|endoftext|>\")\nPREFIX = [SOT, BN, TR, NT]\nassert all(i is not None and i > 0 for i in PREFIX + [EOT]), \"special tokens not found\"\nif GRAD_CKPT:\n    # use_reentrant=False is the documented-safe variant when NO parameter requires\n    # grad and only the input does, which is exactly our case. Verified end to end\n    # below; do not assume it, the failure would be a silent zero gradient.\n    model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={\"use_reentrant\": False})\n    print(\"gradient checkpointing ON (use_reentrant=False)\")\n\nprint(f\"params {sum(p.numel() for p in model.parameters())/1e6:.0f}M | \"\n      f\"prefix {PREFIX} | eot {EOT} | mels {fe.feature_size}\")\n\n# Derive the generation cap from the ACTUAL targets, over all 100 rather than the\n# selected N_EVAL, so switching N_EVAL cannot silently truncate anything.\n_need = [len(tok(norm_bn(t), add_special_tokens=False).input_ids) for t in tg.target_text]\nMAX_NEW = int(max(_need)) + len(PREFIX) + 24\nprint(f\"target length in Whisper tokens: mean {np.mean(_need):.1f}  max {max(_need)}\")\nprint(f\"  ladder char vocab, same targets: mean {tg.target_tokens.mean():.1f}  \"\n      f\"max {tg.target_tokens.max()}  ->  {np.mean(_need)/tg.target_tokens.mean():.2f}x blowup\")\nprint(f\"  MAX_NEW set to {MAX_NEW}\")\nprint(\"  Whisper's BPE is byte-level and Bengali is not in its merge table, so each\")\nprint(\"  target needs about twice as many autoregressive steps to land exactly right.\")\nprint(\"  That is a genuine architectural asymmetry, not a nuisance -- report it.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"48b8ee92-ef91-4a86-a013-c62925feb3bb","cell_type":"markdown","source":"## 5 · Differentiable log-mel front end (trap 1) — with self-tests that must pass","metadata":{}},{"id":"0f301425-ffa8-4229-a4e8-7e5d16496f0e","cell_type":"code","source":"N_SAMPLES, N_FFT, HOP = fe.n_samples, fe.n_fft, fe.hop_length\nMELF = torch.from_numpy(fe.mel_filters).float().to(DEVICE)\nWIN  = torch.hann_window(N_FFT).to(DEVICE)\n\ndef logmel(wav):\n    \"\"\"Whisper log-mel, differentiable end to end. wav: (B, T) float32 in [-1, 1].\n\n    Mirrors transformers' _torch_extract_fbank_features exactly, minus the\n    torch.from_numpy(...) that severs the graph in the public API.\n\n    |z|^2 is written re^2 + im^2, NOT .abs()**2: abs() of a complex tensor has an\n    undefined derivative at the origin, and an all-silence frame sits exactly there.\n    \"\"\"\n    B, T = wav.shape\n    wav = F.pad(wav, (0, N_SAMPLES - T)) if T < N_SAMPLES else wav[:, :N_SAMPLES]\n    st = torch.stft(wav, N_FFT, HOP, window=WIN, return_complex=True,\n                    center=True, pad_mode=\"reflect\")[..., :-1]\n    mel = MELF.T @ (st.real.pow(2) + st.imag.pow(2))\n    lg  = torch.clamp(mel, min=1e-10).log10()\n    lg  = torch.maximum(lg, lg.amax(dim=(1, 2), keepdim=True) - 8.0)\n    return (lg + 4.0) / 4.0\n\ndef feats(x):\n    \"\"\"Log-mel in fp32, cast to the model's dtype at the boundary.\n\n    The STFT, the 1e-10 clamp and the log need fp32 range -- computing them in fp16\n    would break the 1e-4 agreement with HuggingFace that the self-test above asserts.\n    But if the model was loaded with torch_dtype=torch.float16 (which large-v3 needs\n    to fit a T4), feeding it fp32 features raises\n        RuntimeError: Input type (float) and bias type (c10::Half) should be the same\n    So: compute in fp32, cast once, here. The cast is differentiable, so the gradient\n    still reaches the waveform -- verified fp32 grad, all entries non-zero and finite.\n    \"\"\"\n    return logmel(x).to(model.dtype)\n\n# ---- SELF-TEST 1: numerical agreement with HuggingFace -------------------------\n_t = torch.manual_seed(0)\n_w = ((torch.rand(2, SR * 5) * 2 - 1) * 0.5)\n_ref = torch.from_numpy(fe(_w.numpy(), sampling_rate=SR, return_tensors=\"np\").input_features)\n_d = (_ref.to(DEVICE) - logmel(_w.to(DEVICE))).abs().max().item()\nprint(f\"self-test 1  max|mine - HF| = {_d:.3e}\")\nassert _d < 1e-4, (\n    f\"log-mel does not match HuggingFace (max diff {_d:.3e}). Every number from this \"\n    \"notebook would be measured on a different front end than Whisper actually uses.\")\n\n# ---- SELF-TEST 2: the gradient actually reaches the waveform -------------------\n_wg = ((torch.rand(2, SR * 5) * 2 - 1) * 0.5).to(DEVICE).requires_grad_(True)\nlogmel(_wg).sum().backward()\n_g = _wg.grad\n_nz, _nan = int((_g != 0).sum()), int((~torch.isfinite(_g)).sum())\nprint(f\"self-test 2  grad nonzero {_nz}/{_g.numel()}  non-finite {_nan}  \"\n      f\"absmax {_g.abs().max():.3e}\")\nassert _nz > 0 and _nan == 0, (\n    \"Gradient does not reach the waveform, or is non-finite. This is THE failure mode \"\n    \"this task is known for. Do not proceed -- the attack would silently optimise nothing.\")\nprint(\"front end OK -- both self-tests passed\")\ndel _w, _ref, _wg, _g","metadata":{},"outputs":[],"execution_count":null},{"id":"00e75914-d1e0-483b-b3c2-6286ae21d185","cell_type":"markdown","source":"## 6 · Clean WER **gate** — can Whisper read Bengali at all?","metadata":{}},{"id":"ccecff5d-ccf2-4e32-87d3-9b080c55326b","cell_type":"code","source":"ZW = dict.fromkeys(map(ord, \"‌‍\"), None)\ndef norm_bn(s):\n    return re.sub(r\"\\s+\", \" \", unicodedata.normalize(\"NFC\", str(s)).translate(ZW)).strip()\n\ndef wer(ref, hyp):\n    r, h = norm_bn(ref).split(), norm_bn(hyp).split()\n    if not r: return 0.0 if not h else 1.0\n    d = list(range(len(h) + 1))\n    for i in range(1, len(r) + 1):\n        prev, d[0] = d[0], i\n        for j in range(1, len(h) + 1):\n            cur = d[j]\n            d[j] = min(d[j] + 1, d[j-1] + 1, prev + (r[i-1] != h[j-1]))\n            prev = cur\n    return d[len(h)] / len(r)\n\ndef load_wavs(rows):\n    \"\"\"Real-length batch + validity mask. logmel() pads to 30 s internally, so the\n    perturbation never touches the 30 s zero-padding.\"\"\"\n    ws = []\n    for _, r in rows.iterrows():\n        a, _ = librosa.load(r[\"path\"], sr=SR, mono=True)\n        a = np.asarray(a, np.float32)\n        ws.append(a / max(np.abs(a).max(), 1.0))\n    T = max(len(w) for w in ws)\n    x = torch.zeros(len(ws), T); m = torch.zeros(len(ws), T)\n    for i, w in enumerate(ws):\n        x[i, :len(w)] = torch.from_numpy(w); m[i, :len(w)] = 1.0\n    return x.to(DEVICE), m.to(DEVICE)\n\n@torch.no_grad()\ndef transcribe(x):\n    ids = model.generate(input_features=feats(x), language=\"bn\", task=\"transcribe\",\n                         num_beams=1, do_sample=False, max_new_tokens=MAX_NEW)\n    return [norm_bn(s) for s in tok.batch_decode(ids, skip_special_tokens=True)]","metadata":{},"outputs":[],"execution_count":null},{"id":"8c9e6ad7-2b4e-41d4-b4ad-27d84de8d6e6","cell_type":"code","source":"t0 = time.time(); clean_hyp = []\nfor i in range(0, len(ev), BATCH_SIZE):\n    xb, _ = load_wavs(ev.iloc[i:i+BATCH_SIZE])\n    clean_hyp += transcribe(xb)\nev[\"clean_decoded\"] = clean_hyp\nev[\"clean_wer\"] = [wer(r.text, r.clean_decoded) for r in ev.itertuples()]\nCLEAN_WER = float(ev.clean_wer.mean())\nprint(f\"clean transcription of {len(ev)} clips in {time.time()-t0:.0f}s\")\nprint(f\"\\nWHISPER CLEAN WER = {CLEAN_WER:.1%}   exact matches {int((ev.clean_decoded==ev.text.map(norm_bn)).sum())}/{len(ev)}\")\nprint(\"\\nLOOK AT THESE. If they are English, the language pin failed and everything below is void.\\n\")\nfor r in ev.head(4).itertuples():\n    print(f\"  ref : {r.text}\")\n    print(f\"  hyp : {r.clean_decoded}\")\n    print(f\"  wer : {r.clean_wer:.1%}\\n\")\n\nLADDER = {\"1 h final\": 0.6471, \"5 h final\": 0.4755, \"25 h final\": 0.3235,\n          \"100 h final\": 0.2500}\nprint(\"ladder clean WER on the 25-utt spine, for scale:\")\nfor k, v in LADDER.items():\n    print(f\"  {k:12s} {v:.1%}\")\n\n# ---------------------------------------------------------------------------\n# THE GATE IS A SCRIPT GATE, NOT A WER GATE.\n#\n# The earlier 75% clean-WER threshold was the wrong criterion and is removed.\n# Clean accuracy does not bound targeted attack success: PGD optimises directly\n# for the target string, it does not need the model to be accurate. The ladder\n# proves it -- the 1 h rung has clean WER 64.7% on this very set and is 80%\n# attackable at this very budget.\n#\n# What genuinely makes an attack result uninterpretable is the model not emitting\n# the target SCRIPT at all. Asking a model that outputs Devanagari to produce an\n# exact Bengali string measures its script prior, not its adversarial robustness.\n# whisper-small does exactly that: 0% Bengali characters, 64% Devanagari.\n# ---------------------------------------------------------------------------\ndef _script_mix(ss):\n    d = {}\n    for s in ss:\n        for ch in str(s):\n            if   \"\\u0980\" <= ch <= \"\\u09ff\": k = \"Bengali\"\n            elif \"\\u0900\" <= ch <= \"\\u097f\": k = \"Devanagari\"\n            elif \"\\u0600\" <= ch <= \"\\u06ff\": k = \"Arabic\"\n            elif ch.isascii() and ch.isalpha(): k = \"Latin\"\n            elif ch.isalpha(): k = \"other\"\n            else: continue\n            d[k] = d.get(k, 0) + 1\n    tot = sum(d.values()) or 1\n    return {k: v / tot for k, v in sorted(d.items(), key=lambda x: -x[1])}\n\nMIX = _script_mix(ev.clean_decoded)\nBN_FRAC = MIX.get(\"Bengali\", 0.0)\nprint(\"\\noutput script mix: \" + \"  \".join(f\"{k} {100*v:.0f}%\" for k, v in MIX.items()))\nGATE_OK = BN_FRAC >= 0.70\nif not GATE_OK:\n    print(f\"\\n*** GATE FAILED: only {100*BN_FRAC:.0f}% of output characters are Bengali. ***\")\n    print(\"This model does not write Bengali, so a targeted Bengali attack would measure\")\n    print(\"its script prior rather than its robustness. Report the script mix and the clean\")\n    print(\"WER, and STOP. Run NB-11b to pick a model that does write Bengali.\")\nelse:\n    print(f\"\\nGATE PASSED: {100*BN_FRAC:.0f}% Bengali output. clean WER {CLEAN_WER:.1%}, \"\n          f\"CER context from NB-11b.\")\n    print(\"  Clean WER is reported as context, NOT as a pass/fail test -- see the comment\")\n    print(\"  above for why. For scale, the ladder's 1 h rung sits at 64.7% clean WER here\")\n    print(\"  and is 80% attackable at eps=2e-3.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"22ca040e-4669-461a-9961-fc6c968a50fe","cell_type":"markdown","source":"## 7 · Loss and the success criterion\n\nCross-entropy is **teacher-forced**. Success is **not**. A perturbation that lowers teacher-forced\nloss can still fail under free-running generation, because errors compound across decoding steps.\nThe only criterion used below is *`model.generate()` reproduces the target string exactly, after\nthe same `norm_bn` normalisation the ladder uses, on int16-rounded audio*.","metadata":{}},{"id":"43610488-5661-44dc-a96d-a455783e3260","cell_type":"code","source":"def build_labels(targets):\n    \"\"\"Decoder inputs and labels for a batch of target strings.\n\n    full  = [SOT, bn, transcribe, notimestamps] + target_tokens + [EOT]\n    dec_in = full[:-1]   labels = full[1:]\n    The three prefix-predicting positions are masked to -100: those tokens are\n    forced at generation time, so optimising toward them is wasted gradient.\n    \"\"\"\n    seqs = [PREFIX + tok(t, add_special_tokens=False).input_ids + [EOT] for t in targets]\n    L = max(len(s) for s in seqs)\n    di = torch.full((len(seqs), L - 1), EOT, dtype=torch.long)\n    lb = torch.full((len(seqs), L - 1), -100, dtype=torch.long)\n    for i, s in enumerate(seqs):\n        di[i, :len(s)-1] = torch.tensor(s[:-1])\n        lb[i, :len(s)-1] = torch.tensor(s[1:])\n        lb[i, :len(PREFIX)-1] = -100\n    return di.to(DEVICE), lb.to(DEVICE), [len(s) for s in seqs]\n\ndef ce_per_sample(x, di, lb):\n    out = model(input_features=feats(x), decoder_input_ids=di)\n    lp = F.cross_entropy(out.logits.float().transpose(1, 2), lb,\n                         ignore_index=-100, reduction=\"none\")\n    n = (lb != -100).sum(-1).clamp(min=1)\n    return lp.sum(-1) / n\n\ndef int16_rt(x):\n    \"\"\"The ladder rounds to int16 before every verification, so a 'success' is one\n    that survives being written to a real 16-bit WAV. Keep it or the numbers are\n    not comparable.\"\"\"\n    return torch.clamp((x * 32767.0).round(), -32768.0, 32767.0) / 32767.0\n\ndef snr_db(clean, adv, mask):\n    n = (adv - clean) * mask; s = clean * mask\n    return (10 * torch.log10(s.pow(2).sum(-1) / (n.pow(2).sum(-1) + 1e-12))).cpu().numpy()\n\nassert MAX_NEW is not None, \"cell 5 did not run: MAX_NEW is still None\"\n_lab_di, _lab_lb, _lab_len = build_labels(ev.target_text.map(norm_bn).tolist()[:2])\nprint(f\"label builder: dec_in {tuple(_lab_di.shape)}  labels {tuple(_lab_lb.shape)}  \"\n      f\"masked {int((_lab_lb == -100).sum())} positions\")\nassert max(_lab_len) <= MAX_NEW, \"MAX_NEW is below a target length\"","metadata":{},"outputs":[],"execution_count":null},{"id":"e1b9a9fb-fa99-489a-af04-351d3f4552d0","cell_type":"markdown","source":"## 7b · End-to-end gradient self-test — through the **whole model**\n\nCell 6 proved the gradient survives the log-mel. This proves it survives the log-mel *and* the\nencoder *and* the decoder *and* gradient checkpointing, which is the combination that actually\nruns. Checkpointing with `use_reentrant=True` and no parameter requiring grad is a known way to\nget a silent zero gradient, and a silent zero gradient produces a perfectly plausible 0% ASR.","metadata":{}},{"id":"796a3316-3ba9-473e-bcea-f76d3d88f81f","cell_type":"code","source":"_x = ((torch.rand(1, SR * 4) * 2 - 1) * 0.3).to(DEVICE).requires_grad_(True)\n_di, _lb, _ = build_labels([norm_bn(ev.target_text.iloc[0])])\nce_per_sample(_x, _di, _lb).sum().backward()\n_g = _x.grad\n_nz = 0 if _g is None else int((_g != 0).sum())\n_bad = 0 if _g is None else int((~torch.isfinite(_g)).sum())\nprint(f\"end-to-end grad: nonzero {_nz}/{_x.numel()}  non-finite {_bad}  \"\n      f\"absmax {0.0 if _g is None else _g.abs().max():.3e}\")\nassert _nz > 0 and _bad == 0, (\n    \"The gradient does NOT reach the waveform through the full model. If GRAD_CKPT is on, \"\n    \"that is the likely cause -- set GRAD_CKPT = False and rerun this cell. Do not proceed: \"\n    \"a zero gradient produces a completely plausible 0% ASR that means nothing.\")\nprint(\"end-to-end gradient OK\")\ndel _x, _g","metadata":{},"outputs":[],"execution_count":null},{"id":"3fd2909e-6e1a-4d00-83f7-9e1706efbc85","cell_type":"markdown","source":"## 8 · PGD on the waveform","metadata":{}},{"id":"8fe26e34-2c4a-46e4-a5d1-6192304ffd00","cell_type":"code","source":"def pgd(x, mask, targets, eps, steps=None, lr=None, check_every=None, verbose=False):\n    \"\"\"L-inf PGD with Adam, per-sample early exit on verified generation success.\n    Returns (success bool array, delta, first-hit step per sample).\"\"\"\n    steps = steps or ATTACK_STEPS; lr = lr or LR; check_every = check_every or CHECK_EVERY\n    B = x.shape[0]\n    di, lb, _ = build_labels(targets)\n    delta = torch.zeros_like(x, requires_grad=True)\n    opt = torch.optim.Adam([delta], lr=lr)\n    scaler = torch.amp.GradScaler(\"cuda\", enabled=(USE_AMP and DEVICE == \"cuda\"))\n    active = torch.ones(B, dtype=torch.bool, device=DEVICE)\n    frozen = torch.zeros_like(x)\n    success = np.zeros(B, dtype=bool); hit_at = np.full(B, -1)\n    ce0 = None\n\n    for step in range(steps):\n        if not active.any(): break\n        x_adv = torch.clamp(x + delta, -1, 1) * mask\n        with torch.amp.autocast(\"cuda\", enabled=(USE_AMP and DEVICE == \"cuda\")):\n            per = ce_per_sample(x_adv, di, lb)\n        if ce0 is None: ce0 = per.detach().clone()\n        loss = (per * active.float()).sum() / active.float().sum().clamp(min=1)\n        opt.zero_grad(); scaler.scale(loss).backward(); scaler.unscale_(opt)\n        if delta.grad is not None:\n            delta.grad[~active] = 0\n        scaler.step(opt); scaler.update()\n        with torch.no_grad():\n            delta.data = delta.data.clamp(-eps, eps)\n\n        if (step + 1) % check_every == 0:\n            hyp = transcribe(int16_rt(torch.clamp(x + delta, -1, 1)) * mask)\n            for i in range(B):\n                if active[i] and hyp[i] == norm_bn(targets[i]):\n                    frozen[i] = delta[i].detach(); success[i] = True\n                    hit_at[i] = step + 1; active[i] = False\n            if verbose:\n                print(f\"    step {step+1:4d}  CE {loss.item():7.4f}  \"\n                      f\"hit {int(success.sum())}/{B}\")\n    frozen[torch.from_numpy(~success).to(DEVICE)] = delta.detach()[torch.from_numpy(~success).to(DEVICE)]\n    return success, frozen.detach(), hit_at, ce0.cpu().numpy(), per.detach().cpu().numpy()","metadata":{},"outputs":[],"execution_count":null},{"id":"d7a3c295-7370-45d6-bbf7-2612e0ea13be","cell_type":"markdown","source":"## 9A · STAGE = `calibrate`\n\n~10 GPU-minutes on 3 utterances. It answers one question: **is `eps = 2e-3` reachable at all on\nthis architecture, and is `LR` in the right range?** Adam is scale-free per-parameter, so `lr`\nsets how fast `delta` fills its box: at `lr = 1e-4` the box is full in ~20 steps and the\nremaining steps only change signs. If nothing succeeds at the largest eps, the problem is the\noptimiser, not the budget.","metadata":{}},{"id":"5ddfb73c-c3e2-4f1d-948d-b16c6721dbec","cell_type":"code","source":"if STAGE == \"calibrate\":\n    assert GATE_OK, \"clean-WER gate failed -- calibrating an unreadable model is pointless\"\n    sub = ev.head(3)\n    xb, mb = load_wavs(sub); tgts = sub.target_text.map(norm_bn).tolist()\n    grid = [(e, l) for e in [EPS_MAIN] + EPS_EXTRA for l in (1e-4, 1e-3)]\n    print(f\"{'eps':>9s}{'lr':>8s}{'CE start':>10s}{'CE end':>9s}{'hits':>7s}{'sec':>7s}\")\n    cal = []\n    for e, l in grid:\n        t0 = time.time()\n        s, d, ha, c0, c1 = pgd(xb, mb, tgts, e, steps=200, lr=l, check_every=50)\n        dt = time.time() - t0\n        cal.append(dict(eps=e, lr=l, ce_start=float(c0.mean()), ce_end=float(c1.mean()),\n                        hits=int(s.sum()), n=len(sub), sec=round(dt, 1)))\n        print(f\"{e:9.1e}{l:8.1e}{c0.mean():10.4f}{c1.mean():9.4f}\"\n              f\"{int(s.sum()):5d}/{len(sub)}{dt:7.0f}\")\n    cal = pd.DataFrame(cal)\n    per_step = cal.sec.sum() / (len(grid) * 200)\n    print(f\"\\nmean {per_step*1000:.0f} ms per PGD step at batch {len(sub)}\")\n    print(f\"projected full run: {len(ev)} utts x {ATTACK_STEPS} steps, batch {BATCH_SIZE}\"\n          f\"  ~= {per_step*ATTACK_STEPS*len(ev)/BATCH_SIZE/3600:.1f} GPU-h per eps value,\"\n          f\"  x{1+len(EPS_EXTRA)} eps values\")\n    print(f\"  NOTE: that figure already includes a generate() verification every \"\n          f\"{CHECK_EVERY} steps, and generation of up to {MAX_NEW} tokens is the\")\n    print(\"  dominant cost, not the gradient step. If the projection does not fit your\")\n    print(\"  quota, raise CHECK_EVERY before you touch ATTACK_STEPS -- checking less\")\n    print(\"  often costs you resolution on hit_step, not attack strength.\")\n    best = cal.sort_values([\"hits\", \"ce_end\"], ascending=[False, True]).iloc[0]\n    print(\"\\n--- read this ---\")\n    if cal.hits.sum() == 0:\n        print(\"NOTHING succeeded, at any budget. Do NOT run the attack stage.\")\n        print(\"Either the attack does not work here, or CE is not descending -- check the\")\n        print(\"CE start/end columns. If CE barely moved even at the largest eps, raise LR\")\n        print(\"by 10x and re-calibrate once. If it still does not move, stop and report it.\")\n    elif int(cal[cal.eps == EPS_MAIN].hits.sum()) == 0:\n        print(f\"Nothing succeeded at the ladder's eps={EPS_MAIN:.0e}, but larger budgets did.\")\n        print(\"That is itself the result: Whisper needs more perturbation than our rungs.\")\n        print(f\"Run the attack stage -- ASR@{EPS_MAIN:.0e} = 0% is a reportable number, and\")\n        print(\"EPS_EXTRA gives the second datapoint that makes it interpretable.\")\n    else:\n        print(f\"eps={EPS_MAIN:.0e} is reachable. Set LR = {best.lr:.0e} if it differs from\")\n        print(f\"the current {LR:.0e}, then set STAGE = 'attack'.\")\n    cal.to_csv(\"/kaggle/working/whisper_calibration.csv\", index=False)","metadata":{},"outputs":[],"execution_count":null},{"id":"711738da-fbcf-4961-93e6-4084f7d289df","cell_type":"markdown","source":"## 9B · STAGE = `attack` — the headline number","metadata":{}},{"id":"1e9c7be3-c7ad-4477-b2b9-7c3ad7b59f06","cell_type":"code","source":"if STAGE == \"attack\":\n    assert GATE_OK, \"clean-WER gate failed -- do not produce an attack number\"\n    EPS_LIST = [EPS_MAIN] + EPS_EXTRA\n    rows, t_start = [], time.time()\n    for eps in EPS_LIST:\n        hits = 0\n        print(f\"\\n=== eps = {eps:.1e} ===\")\n        for i in range(0, len(ev), BATCH_SIZE):\n            sub = ev.iloc[i:i+BATCH_SIZE]\n            xb, mb = load_wavs(sub); tgts = sub.target_text.map(norm_bn).tolist()\n            s, d, ha, c0, c1 = pgd(xb, mb, tgts, eps)\n            xa = int16_rt(torch.clamp(xb + d, -1, 1)) * mb\n            dec = transcribe(xa); sn = snr_db(xb, xa, mb)\n            for j, (_, r) in enumerate(sub.iterrows()):\n                rows.append(dict(utt_id=r.utt_id, model=MODEL_ID, arch=\"whisper-enc-dec\",\n                                 edit_type=EDIT_TYPE, eps=eps, duration=r.duration,\n                                 source_text=r.text, target_text=r.target_text,\n                                 target_tokens_whisper=len(tok(norm_bn(r.target_text),\n                                     add_special_tokens=False).input_ids),\n                                 target_tokens_ladder=r.target_tokens,\n                                 success=bool(s[j]), hit_step=int(ha[j]),\n                                 ce_start=float(c0[j]), ce_end=float(c1[j]),\n                                 snr_db=float(sn[j]), decoded=dec[j],\n                                 clean_decoded=r.clean_decoded, clean_wer=r.clean_wer,\n                                 exact_match=bool(dec[j] == norm_bn(r.target_text))))\n            hits += int(s.sum())\n            print(f\"  {i+len(sub):3d}/{len(ev)}  cumulative hits {hits}  \"\n                  f\"({time.time()-t_start:.0f}s)\")\n        print(f\"  ASR@{eps:.1e} = {100*hits/len(ev):.1f}%  ({hits}/{len(ev)})\")\n    res = pd.DataFrame(rows)\n    res.to_csv(\"/kaggle/working/whisper_attack_results.csv\", index=False)\n    print(f\"\\ntotal {(time.time()-t_start)/60:.1f} min, {len(res)} rows written\")","metadata":{},"outputs":[],"execution_count":null},{"id":"bbe5f21e-5316-4888-8754-cdbf60917d98","cell_type":"markdown","source":"## 10 · STAGE = `transfer` — free, the adversarial WAVs already exist","metadata":{}},{"id":"adc289aa-37e1-4cca-8c83-f6dfc07fb4a5","cell_type":"code","source":"if STAGE == \"transfer\":\n    # NB-07 wrote every successful adversarial waveform to pipeline-survival-rung*/wav/\n    # with SAVE_WAVS=True precisely so this cost would be paid once. No attack compute.\n    WAVD = next((f\"{p}/wav\" for p in inputs if os.path.isdir(f\"{p}/wav\")\n                 and \"pipeline-survival\" in p), None)\n    assert WAVD, (\"pipeline-survival-rung100 not attached (looking for its wav/ dir). \"\n                  f\"Visible: {[os.path.basename(p) for p in inputs][:20]}\")\n    wavs = sorted(glob.glob(f\"{WAVD}/*adv*.wav\")) or sorted(glob.glob(f\"{WAVD}/*.wav\"))\n    assert wavs, f\"no wav files under {WAVD}\"\n    print(f\"{len(wavs)} waveforms from {WAVD}\")\n\n    tmap = {r.utt_id: norm_bn(r.target_text) for r in tg.itertuples()}\n    smap = {r.utt_id: norm_bn(r.text) for r in tg.itertuples()}\n    trows = []\n    for i in range(0, len(wavs), BATCH_SIZE):\n        chunk = wavs[i:i+BATCH_SIZE]; arrs = []\n        for w in chunk:\n            a, _ = librosa.load(w, sr=SR, mono=True)\n            a = np.asarray(a, np.float32); arrs.append(a / max(np.abs(a).max(), 1.0))\n        T = max(len(a) for a in arrs)\n        x = torch.zeros(len(arrs), T); m = torch.zeros(len(arrs), T)\n        for j, a in enumerate(arrs):\n            x[j, :len(a)] = torch.from_numpy(a); m[j, :len(a)] = 1.0\n        dec = transcribe(x.to(DEVICE))\n        for j, w in enumerate(chunk):\n            uid = os.path.basename(w).split(\"_\")[0].split(\".\")[0]\n            trows.append(dict(wav=os.path.basename(w), utt_id=uid, decoded=dec[j],\n                              target=tmap.get(uid, \"\"), source=smap.get(uid, \"\"),\n                              hits_target=dec[j] == tmap.get(uid, \"\\x00\"),\n                              returns_source=dec[j] == smap.get(uid, \"\\x00\")))\n    tr = pd.DataFrame(trows)\n    tr.to_csv(\"/kaggle/working/whisper_transfer.csv\", index=False)\n    known = tr[tr.target != \"\"]\n    print(f\"\\nmatched {len(known)}/{len(tr)} filenames to a known utt_id\")\n    if len(known):\n        print(f\"  reproduces the XLS-R attacker's target : {int(known.hits_target.sum())}/{len(known)}\")\n        print(f\"  returns the true transcript instead    : {int(known.returns_source.sum())}/{len(known)}\")\n        print(\"\\n  Targeted transfer across architectures is expected to be ~0. A zero here is\")\n        print(\"  a clean, reportable negative: the perturbation is specific to the CTC model's\")\n        print(\"  decision surface, not a property of the audio.\")\n    else:\n        print(\"  Filenames did not map to utt_ids. Print a few and fix the parse:\")\n        print(\"   \", [os.path.basename(w) for w in wavs[:5]])","metadata":{},"outputs":[],"execution_count":null},{"id":"775ee6ac-7bf5-47af-b845-46802a999007","cell_type":"markdown","source":"## 11 · Results, provenance, and what to do with them","metadata":{}},{"id":"d0cae391-a51a-4e28-8ca9-0b59e2f30010","cell_type":"code","source":"PROV = dict(notebook=\"NB-11-whisper-arm\", stage=STAGE, model=MODEL_ID,\n            n_eval=int(N_EVAL), edit_type=EDIT_TYPE, eps_main=EPS_MAIN,\n            eps_extra=EPS_EXTRA, seed=SEED, steps=ATTACK_STEPS, lr=LR,\n            batch_size=BATCH_SIZE, use_amp=USE_AMP, device=DEVICE,\n            torch=torch.__version__, transformers=transformers.__version__,\n            targets_sha256=TGT_SHA, clean_wer=round(CLEAN_WER, 4),\n            gate_passed=bool(GATE_OK),\n            timestamp=time.strftime(\"%Y-%m-%d %H:%M:%S\"))\n\nif STAGE == \"attack\":\n    summ = (res.groupby(\"eps\")\n              .agg(n=(\"success\", \"size\"), hits=(\"success\", \"sum\"),\n                   median_snr=(\"snr_db\", \"median\"))\n              .assign(asr=lambda d: 100 * d.hits / d.n).reset_index())\n    PROV[\"asr_by_eps\"] = {f\"{r.eps:.1e}\": round(float(r.asr), 1) for r in summ.itertuples()}\n    print(\"=\" * 66); print(f\"WHISPER ARM — {MODEL_ID}, n = {len(ev)}\"); print(\"=\" * 66)\n    print(f\"clean WER                {CLEAN_WER:.1%}\")\n    for r in summ.itertuples():\n        print(f\"ASR @ eps {r.eps:.1e}     {r.asr:5.1f}%   ({int(r.hits)}/{int(r.n)})\"\n              f\"   median SNR {r.median_snr:.1f} dB\")\n    print(\"\\ncompare with the ladder at the SAME eps=2e-3, same targets, n=100:\")\n    for k, v in [(\"1 h final\", 75), (\"5 h final\", 79), (\"25 h final\", 66),\n                 (\"100 h ms-1000\", 37), (\"100 h ms-2000\", 46), (\"100 h ms-3500\", 53)]:\n        print(f\"  {k:15s} {v:3d}%\")\n    print(\"\\nHOW TO WRITE IT. Whisper is not a rung. Its pretraining data is not ours and\")\n    print(\"is not controlled, so this does NOT extend the dose-response curve. It answers\")\n    print(\"exactly one question -- whether the fragility is specific to CTC -- and the\")\n    print(\"sentence in the paper must carry that caveat in the same breath as the number.\")\n\njson.dump(PROV, open(\"/kaggle/working/_provenance.json\", \"w\"), ensure_ascii=False, indent=2)\nprint(\"\\n\" + json.dumps(PROV, ensure_ascii=False, indent=2))\nprint(\"\\nSave version -> publish /kaggle/working as `whisper-arm-results`.\")\nprint(\"CHECK THE OUTPUT FILE LIST BEFORE PUBLISHING: publishing replaces the dataset and\")\nprint(\"deletes anything a partial run did not regenerate.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"259a6bfc-7041-469e-84b8-813413a30815","cell_type":"markdown","source":"---\n\n## Order of operations\n\n1. `STAGE = \"calibrate\"` — ~10 GPU-min. Read the verdict it prints.\n2. `STAGE = \"attack\"` — 1–3 GPU-h depending on `N_EVAL`. Produces ASR@2e-3.\n3. `STAGE = \"transfer\"` — minutes, no attack compute. Needs `pipeline-survival-rung100`.\n\nRecord each in `RUNLOG.md` with the settings block, the clean WER, and the ASR table.\n\n## Every outcome here is reportable\n\n| outcome | what it means | paper sentence |\n|---|---|---|\n| clean WER > 75% | Whisper cannot read this Bengali | *\"Off-the-shelf `whisper-small` reaches only X% WER on our Bengali test set, so we could not evaluate it as a second architecture.\"* |\n| ASR@2e-3 ≈ 0, larger eps works | Whisper costs more to attack | *\"At the ladder's budget the attack does not transfer to an autoregressive model; it requires Nx more perturbation.\"* |\n| ASR@2e-3 comparable to the rungs | fragility is not CTC-specific | *\"A second architecture with far more pretraining shows comparable targeted attack cost.\"* |\n| nothing succeeds at any eps | the attack, not the model | Report where it stopped. **Do not report a 0% ASR you cannot distinguish from a broken attack** — that is what the calibration stage exists to separate. |\n\nThe single-architecture limitation is already stated in the paper. This task **upgrades** that\nlimitation; it does not rescue anything, and it is allowed to fail.","metadata":{}}]}