{"cells": [{"cell_type": "markdown", "id": "d186bdc0", "metadata": {}, "source": "# How much does hand-shape alone carry the signal?\n\n*A head-to-head of a hand-only single-frame classifier vs a small temporal model, on a signer-holdout split \u2014 to find out how much of the signal is really in the shape of the hands at one moment vs. how they move over time.*\n\n**Question.** How much does hand-shape prior dominate temporal information on isolated signs in Google ISLR?\n\n**Hypothesis.** A hand-only single-frame MLP reaches within 10 percentage points of a small 1D-conv temporal model on signer-holdout test accuracy. If true, hand-shape priors carry most of the signal on this dataset and the temporal value-add is a smaller bet than the field assumes.\n\n**Dataset.** Google Isolated Sign Language Recognition (`asl-signs`), 250 signs \u00d7 21 signers \u00d7 ~94k clips. License: Kaggle competition terms.\n\n**Split.** **Signer-holdout 17/2/2** (ADR 0003, seed 42) \u2014 two signers held out for val, two for test, no signer overlap between train and eval. This departs from every published Kaggle winner's random split; see Parley Notebook 00 \u00a77 for why.\n\n**Baselines reported head-to-head with each model:** random chance (1/250 = 0.4%), majority class, linear probe on the same hand features.\n\n**Statistical floor.** Mean \u00b1 std over 3 seeds (auto-escalates to 5 if std/mean > 0.05). 1-hour wall-clock cap per run.\n\n**How to reproduce.** The 6 training runs (3 seeds \u00d7 2 models) ran offline via `scripts/run_notebook_01_training.py` and produced the aggregation parquets attached as `parley-nb01-caches`. This notebook reads those caches \u2014 no training happens here, so Kaggle runs complete in ~1 minute. To re-train end-to-end, clone the Parley repo and run the orchestrator script locally.\n"}, {"cell_type": "code", "execution_count": 1, "id": "a6fbb830", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:44.185009Z", "iopub.status.busy": "2026-04-23T02:42:44.184729Z", "iopub.status.idle": "2026-04-23T02:42:46.908060Z", "shell.execute_reply": "2026-04-23T02:42:46.907683Z"}}, "outputs": [], "source": "%matplotlib inline\nfrom __future__ import annotations\n\nfrom dataclasses import dataclass, field\nfrom pathlib import Path\nfrom typing import Callable, Literal\nimport copy\nimport json\nimport subprocess\nimport time\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport torch\nimport torch.nn as nn\n\n# ---------------------------------------------------------------------------\n# Path detection \u2014 Kaggle runtime vs local repo.\n# ---------------------------------------------------------------------------\n_KAGGLE_INPUT = Path(\"/kaggle/input\")\nIN_KAGGLE = _KAGGLE_INPUT.exists()\n\nif IN_KAGGLE:\n    # Kaggle mounts can be flat (/kaggle/input/<slug>/) or nested under\n    # datasets/ and competitions/ \u2014 depends on how the kernel was submitted.\n    # Try both.\n    _ds_candidates = [\n        _KAGGLE_INPUT / \"asl-signs\",\n        _KAGGLE_INPUT / \"competitions\" / \"asl-signs\",\n    ]\n    _ca_candidates = [\n        _KAGGLE_INPUT / \"parley-nb01-caches\",\n        _KAGGLE_INPUT / \"datasets\" / \"truepathventures\" / \"parley-nb01-caches\",\n    ]\n    DATASET_ROOT = next((p for p in _ds_candidates if p.exists()), _ds_candidates[0])\n    CACHE_ROOT = next((p for p in _ca_candidates if p.exists()), _ca_candidates[0])\nelse:\n    try:\n        NB_DIR = Path(__file__).resolve().parent  # type: ignore[name-defined]\n    except NameError:\n        NB_DIR = Path.cwd() if Path.cwd().name == \"notebooks\" else Path.cwd() / \"notebooks\"\n    REPO_ROOT = NB_DIR.parent\n    DATASET_ROOT = REPO_ROOT / \"data\" / \"raw\" / \"google-asl-signs\"\n    CACHE_ROOT = REPO_ROOT / \"data\" / \"processed\" / \"notebook-01\"\n\nSEED = 42\nMANIFEST = \"train.csv\"\n\nsns.set_theme(style=\"whitegrid\", context=\"notebook\")\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\n\nprint(\"Running on:\", \"Kaggle\" if IN_KAGGLE else \"local repo\")\n"}, {"cell_type": "code", "execution_count": 2, "id": "1500a5cc", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:46.909385Z", "iopub.status.busy": "2026-04-23T02:42:46.909270Z", "iopub.status.idle": "2026-04-23T02:42:47.230883Z", "shell.execute_reply": "2026-04-23T02:42:47.230473Z"}}, "outputs": [], "source": "# =========================================================================\n# Inlined from Parley's `signlang/` package (signlang.data, signlang.eval,\n# signlang.models, signlang.training, signlang.plotting) so this notebook\n# runs standalone on Kaggle. The repo keeps each module split out with\n# pytest coverage \u2014 same logic, no differences.\n# =========================================================================\n\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import f1_score\nfrom torch.utils.data import DataLoader\nimport random\nimport hashlib\n\n\n# ---------------------------------------------------------------------------\n# signlang/data/landmarks.py\n# ---------------------------------------------------------------------------\n\n@dataclass(frozen=True)\nclass ClipMetadata:\n    \"\"\"Lightweight metadata for one ISLR clip.\"\"\"\n\n    clip_id: str\n    sign: str\n    signer_id: int\n    frame_count: int\n    missing_rate: float\n\n\n# Canonical landmark order used across the whole project. Any code that\n# indexes into the 543 axis of a loaded clip relies on this ordering.\nLANDMARK_TYPE_COUNTS: dict[str, int] = {\n    \"face\": 468,\n    \"left_hand\": 21,\n    \"pose\": 33,\n    \"right_hand\": 21,\n}\nLANDMARK_ORDER: tuple[str, ...] = (\"face\", \"left_hand\", \"pose\", \"right_hand\")\nN_LANDMARKS: int = sum(LANDMARK_TYPE_COUNTS.values())  # 543\n\n\ndef load_clip(path) -> np.ndarray:\n    \"\"\"Load one ISLR parquet file as a (n_frames, 543, 3) float32 array.\n\n    The second axis is ordered by LANDMARK_ORDER, each block sized by\n    LANDMARK_TYPE_COUNTS. Missing detections become NaN (ISLR stores them as\n    NaN already; any zero-filled rows are preserved as zero, not remapped).\n    \"\"\"\n    df = pd.read_parquet(path)\n    frames = np.sort(df[\"frame\"].unique())\n    n_frames = len(frames)\n    frame_to_row = {f: i for i, f in enumerate(frames)}\n\n    out = np.full((n_frames, N_LANDMARKS, 3), np.nan, dtype=np.float32)\n\n    offsets: dict[str, int] = {}\n    cursor = 0\n    for lm_type in LANDMARK_ORDER:\n        offsets[lm_type] = cursor\n        cursor += LANDMARK_TYPE_COUNTS[lm_type]\n\n    for lm_type in LANDMARK_TYPE_COUNTS:\n        sub = df[df[\"type\"] == lm_type]\n        if sub.empty:\n            continue\n        rows = np.array([frame_to_row[f] for f in sub[\"frame\"].to_numpy()])\n        cols = offsets[lm_type] + sub[\"landmark_index\"].to_numpy(dtype=np.int64)\n        out[rows, cols, 0] = sub[\"x\"].to_numpy(dtype=np.float32)\n        out[rows, cols, 1] = sub[\"y\"].to_numpy(dtype=np.float32)\n        out[rows, cols, 2] = sub[\"z\"].to_numpy(dtype=np.float32)\n\n    return out\n\n\ndef missing_mask(clip: np.ndarray) -> np.ndarray:\n    \"\"\"Return a (n_frames, 543) bool mask where True = missing detection.\"\"\"\n    if clip.ndim != 3 or clip.shape[1] != N_LANDMARKS or clip.shape[2] != 3:\n        raise ValueError(\n            f\"expected clip shape (n_frames, {N_LANDMARKS}, 3), got {clip.shape}\"\n        )\n    nan_missing = np.isnan(clip).any(axis=-1)\n    zero_missing = (clip == 0.0).all(axis=-1)\n    return nan_missing | zero_missing\n\n\ndef list_clips(root, *, manifest_name: str = \"train.csv\") -> list:\n    \"\"\"List clips under a dataset root. Returns one ClipMetadata per manifest row.\"\"\"\n    root = Path(root)\n    manifest = pd.read_csv(root / manifest_name)\n    out = []\n    for _, row in manifest.iterrows():\n        clip = load_clip(root / row[\"path\"])\n        n_frames = clip.shape[0]\n        missing = float(missing_mask(clip).mean()) if n_frames > 0 else 0.0\n        out.append(\n            ClipMetadata(\n                clip_id=str(row[\"sequence_id\"]),\n                sign=str(row[\"sign\"]),\n                signer_id=int(row[\"participant_id\"]),\n                frame_count=int(n_frames),\n                missing_rate=float(missing),\n            )\n        )\n    return out\n\n\n# ---------------------------------------------------------------------------\n# signlang/data/version.py\n# ---------------------------------------------------------------------------\n\ndef dataset_version(root) -> str:\n    \"\"\"Return a 12-char hex digest over (filename, size) for all parquet files.\"\"\"\n    root = Path(root)\n    files = sorted(root.rglob(\"*.parquet\"))\n    manifest = \"\\n\".join(\n        f\"{p.relative_to(root)}|{p.stat().st_size}\" for p in files\n    ).encode()\n    return hashlib.sha256(manifest).hexdigest()[:12]\n\n\n# ---------------------------------------------------------------------------\n# signlang/data/splits.py\n# ---------------------------------------------------------------------------\n\n@dataclass(frozen=True)\nclass SignerHoldoutSplit:\n    train: list[int]\n    val: list[int]\n    test: list[int]\n\n\n@dataclass(frozen=True)\nclass RandomSplit:\n    train: list[int]\n    val: list[int]\n    test: list[int]\n\n\ndef signer_holdout_split(signer_ids, *, seed: int = 42) -> SignerHoldoutSplit:\n    \"\"\"Pick 2 val + 2 test signers deterministically. Remainder is train.\"\"\"\n    sorted_ids = sorted(set(signer_ids))\n    rng = random.Random(seed)\n    picked = rng.sample(sorted_ids, k=4)\n    val = sorted(picked[:2])\n    test = sorted(picked[2:])\n    train = sorted(set(sorted_ids) - set(val) - set(test))\n    return SignerHoldoutSplit(train=train, val=val, test=test)\n\n\ndef apply_split(clips, split):\n    \"\"\"Partition clips by signer_id membership in each set.\"\"\"\n    train_set, val_set, test_set = set(split.train), set(split.val), set(split.test)\n    train = [c for c in clips if c.signer_id in train_set]\n    val = [c for c in clips if c.signer_id in val_set]\n    test = [c for c in clips if c.signer_id in test_set]\n    return train, val, test\n\n\ndef random_split(n: int, *, val_frac: float = 0.1, test_frac: float = 0.1, seed: int = 42) -> RandomSplit:\n    \"\"\"Random split of index range [0, n). Returned as three disjoint index lists.\"\"\"\n    idx = list(range(n))\n    rng = random.Random(seed)\n    rng.shuffle(idx)\n    n_val = int(n * val_frac)\n    n_test = int(n * test_frac)\n    val = sorted(idx[:n_val])\n    test = sorted(idx[n_val : n_val + n_test])\n    train = sorted(idx[n_val + n_test :])\n    return RandomSplit(train=train, val=val, test=test)\n\n\n# ---------------------------------------------------------------------------\n# signlang/data/features.py\n# ---------------------------------------------------------------------------\n\ndef _hand_block_ranges() -> tuple[tuple[int, int], tuple[int, int]]:\n    \"\"\"Return ((left_start, left_end), (right_start, right_end)).\n\n    Hand blocks are NOT contiguous in LANDMARK_ORDER \u2014 pose (33 landmarks)\n    sits between left_hand and right_hand \u2014 so we return both ranges\n    separately and concatenate in the caller.\n    \"\"\"\n    offsets = {}\n    cursor = 0\n    for lm_type in LANDMARK_ORDER:\n        offsets[lm_type] = cursor\n        cursor += LANDMARK_TYPE_COUNTS[lm_type]\n    left_start = offsets[\"left_hand\"]\n    left_end = left_start + LANDMARK_TYPE_COUNTS[\"left_hand\"]\n    right_start = offsets[\"right_hand\"]\n    right_end = right_start + LANDMARK_TYPE_COUNTS[\"right_hand\"]\n    return (left_start, left_end), (right_start, right_end)\n\n\ndef hand_feature_vector(clip: np.ndarray) -> np.ndarray:\n    \"\"\"Return a 126-dim float32 vector of BOTH hands' (x, y, z) at the middle frame.\n\n    Output layout: [left_hand (21 \u00d7 3), right_hand (21 \u00d7 3)] flattened,\n    so v[0:63] is left_hand and v[63:126] is right_hand.\n\n    Falls back to mean-over-frames of each hand block if the middle frame's\n    hand landmarks are entirely NaN. Remaining NaNs become 0.0.\n    \"\"\"\n    if clip.ndim != 3 or clip.shape[1] != 543 or clip.shape[2] != 3:\n        raise ValueError(f\"expected (n_frames, 543, 3), got {clip.shape}\")\n    (ls, le), (rs, re) = _hand_block_ranges()\n    n_frames = clip.shape[0]\n    middle_idx = n_frames // 2\n\n    middle_left = clip[middle_idx, ls:le, :]\n    middle_right = clip[middle_idx, rs:re, :]\n    middle_block = np.concatenate([middle_left, middle_right], axis=0)\n\n    if np.isnan(middle_block).all():\n        mean_left = np.nanmean(clip[:, ls:le, :], axis=0)\n        mean_right = np.nanmean(clip[:, rs:re, :], axis=0)\n        block = np.concatenate([mean_left, mean_right], axis=0)\n    else:\n        block = middle_block\n\n    block = np.nan_to_num(block, nan=0.0)\n    return block.reshape(-1).astype(np.float32)\n\n\n# ---------------------------------------------------------------------------\n# signlang/data/sequences.py\n# ---------------------------------------------------------------------------\n\ndef build_sequence_tensor(\n    clip: np.ndarray,\n    *,\n    target_length: int = 22,\n    zscore_xy: bool = False,\n) -> np.ndarray:\n    \"\"\"Return (target_length, 543, 3) float32, padded/truncated, NaN->0.\n\n    Pad: replicate last frame. Truncate: keep first `target_length` frames.\n    If zscore_xy=True, standardize x and y per clip (mean=0, std=1); z\n    passes through unnormalized.\n    \"\"\"\n    if clip.ndim != 3 or clip.shape[1] != 543 or clip.shape[2] != 3:\n        raise ValueError(f\"expected (n_frames, 543, 3), got {clip.shape}\")\n\n    clip = np.nan_to_num(clip, nan=0.0)\n    n_frames = clip.shape[0]\n\n    if n_frames >= target_length:\n        out = clip[:target_length].copy()\n    else:\n        pad_frames = target_length - n_frames\n        last = clip[-1:].repeat(pad_frames, axis=0) if n_frames > 0 else np.zeros(\n            (pad_frames, 543, 3), dtype=clip.dtype\n        )\n        out = np.concatenate([clip, last], axis=0)\n\n    if zscore_xy:\n        for coord in (0, 1):\n            vals = out[:, :, coord]\n            std = vals.std() or 1.0\n            out[:, :, coord] = (vals - vals.mean()) / std\n\n    return out.astype(np.float32)\n\n\n# ---------------------------------------------------------------------------\n# signlang/eval/baselines.py\n# ---------------------------------------------------------------------------\n\ndef random_baseline_accuracy(*, n_classes: int, n_samples: int = 10_000, seed: int = 42) -> float:\n    \"\"\"Empirical random-baseline accuracy = Monte-Carlo estimate of 1/n_classes.\"\"\"\n    rng = np.random.default_rng(seed)\n    y_true = rng.integers(0, n_classes, n_samples)\n    y_pred = rng.integers(0, n_classes, n_samples)\n    return float((y_true == y_pred).mean())\n\n\ndef majority_class_accuracy(*, y_train: np.ndarray, y_test: np.ndarray) -> float:\n    \"\"\"Predict the most common train class for every test point.\"\"\"\n    vals, counts = np.unique(y_train, return_counts=True)\n    majority = vals[counts.argmax()]\n    return float((y_test == majority).mean())\n\n\ndef linear_probe_accuracy(\n    *, x_train: np.ndarray, y_train: np.ndarray,\n    x_test: np.ndarray, y_test: np.ndarray, seed: int = 42,\n) -> float:\n    \"\"\"Logistic-regression probe on the given features.\"\"\"\n    clf = LogisticRegression(max_iter=500, random_state=seed)\n    clf.fit(x_train, y_train)\n    return float(clf.score(x_test, y_test))\n\n\n# ---------------------------------------------------------------------------\n# signlang/eval/metrics.py\n# ---------------------------------------------------------------------------\n\ndef top_k_accuracy(logits: np.ndarray, y_true: np.ndarray, *, k: int = 1) -> float:\n    top_k = np.argsort(logits, axis=-1)[:, -k:]\n    correct = np.any(top_k == y_true[:, None], axis=1)\n    return float(correct.mean())\n\n\ndef macro_f1(y_true: np.ndarray, y_pred: np.ndarray) -> float:\n    return float(f1_score(y_true, y_pred, average=\"macro\", zero_division=0))\n\n\ndef aggregate_over_seeds(\n    per_seed: list[dict], *, escalate_threshold: float = 0.05,\n) -> pd.DataFrame:\n    \"\"\"Collapse a list of per-seed metric dicts into a long-format (metric, mean, std) frame.\n\n    Adds a `needs_escalation` column: True where std/|mean| > escalate_threshold.\n    \"\"\"\n    if not per_seed:\n        return pd.DataFrame(columns=[\"metric\", \"mean\", \"std\", \"needs_escalation\"])\n    keys = set().union(*(s.keys() for s in per_seed))\n    rows = []\n    for k in sorted(keys):\n        vals = np.array([s.get(k, np.nan) for s in per_seed], dtype=float)\n        mean = float(np.nanmean(vals))\n        std = float(np.nanstd(vals, ddof=1)) if len(vals) > 1 else 0.0\n        rows.append({\n            \"metric\": k, \"mean\": mean, \"std\": std,\n            \"needs_escalation\": std / max(abs(mean), 1e-9) > escalate_threshold,\n        })\n    return pd.DataFrame(rows)\n\n\n# ---------------------------------------------------------------------------\n# signlang/eval/confusion.py\n# ---------------------------------------------------------------------------\n\ndef confusion_matrix(\n    y_true: np.ndarray, y_pred: np.ndarray, *, n_classes: int,\n) -> np.ndarray:\n    \"\"\"Build a confusion matrix from true and predicted labels.\n\n    Returns (n_classes, n_classes) where cm[i, j] = count of samples\n    with true label i and predicted label j.\n    \"\"\"\n    cm = np.zeros((n_classes, n_classes), dtype=np.int64)\n    for t, p in zip(y_true, y_pred):\n        cm[int(t), int(p)] += 1\n    return cm\n\n\ndef top_n_confused_pairs(\n    y_true: np.ndarray, y_pred: np.ndarray, *, n: int, labels: list[str],\n) -> pd.DataFrame:\n    \"\"\"Return top-N (true, predicted) pairs by count, excluding the diagonal.\"\"\"\n    cm = confusion_matrix(y_true, y_pred, n_classes=len(labels))\n    rows = []\n    for i in range(len(labels)):\n        for j in range(len(labels)):\n            if i != j and cm[i, j] > 0:\n                rows.append({\"true\": labels[i], \"predicted\": labels[j], \"count\": int(cm[i, j])})\n    df = pd.DataFrame(rows)\n    if df.empty:\n        return df\n    return df.sort_values(\"count\", ascending=False).head(n).reset_index(drop=True)\n\n\n# ---------------------------------------------------------------------------\n# signlang/eval/failure_gallery.py\n# ---------------------------------------------------------------------------\n\ndef sample_failures(predictions: pd.DataFrame, *, n: int, seed: int = 42) -> pd.DataFrame:\n    \"\"\"Return up to n misclassifications stratified across predicted classes.\n\n    Required columns: clip_id, true, pred.\n    \"\"\"\n    misses = predictions[predictions[\"true\"] != predictions[\"pred\"]].copy()\n    if misses.empty:\n        return misses\n\n    groups = [g for _, g in misses.groupby(\"pred\")]\n    groups_sorted = sorted(groups, key=len, reverse=True)\n    picks: list[pd.DataFrame] = []\n    remaining = n\n\n    for g in groups_sorted:\n        if remaining <= 0:\n            break\n        picks.append(g.sample(1, random_state=seed))\n        remaining -= 1\n\n    if remaining > 0:\n        already = pd.concat(picks).index if picks else pd.Index([])\n        pool = misses.drop(index=already)\n        if not pool.empty:\n            more = pool.sample(n=min(remaining, len(pool)), random_state=seed)\n            picks.append(more)\n\n    return pd.concat(picks).reset_index(drop=True)\n\n\n# ---------------------------------------------------------------------------\n# signlang/models/hand_shape_baseline.py\n# ---------------------------------------------------------------------------\n\nclass HandShapeMLP(nn.Module):\n    \"\"\"Hand-only single-frame MLP baseline.\n\n    Input:  (batch, 126) \u2014 both hands x (x, y, z) at a single frame.\n    Output: (batch, n_classes) \u2014 logits.\n    \"\"\"\n\n    def __init__(self, input_dim: int = 126, n_classes: int = 250, hidden: int = 512):\n        super().__init__()\n        self.net = nn.Sequential(\n            nn.Linear(input_dim, hidden),\n            nn.ReLU(),\n            nn.Dropout(0.3),\n            nn.Linear(hidden, hidden),\n            nn.ReLU(),\n            nn.Dropout(0.3),\n            nn.Linear(hidden, n_classes),\n        )\n\n    def forward(self, x):\n        return self.net(x)\n\n\n# ---------------------------------------------------------------------------\n# signlang/models/temporal_baseline.py\n# ---------------------------------------------------------------------------\n\nclass _ConvBlock(nn.Module):\n    def __init__(self, in_ch: int, out_ch: int, k: int = 5):\n        super().__init__()\n        self.net = nn.Sequential(\n            nn.Conv1d(in_ch, out_ch, kernel_size=k, padding=k // 2),\n            nn.BatchNorm1d(out_ch),\n            nn.ReLU(),\n        )\n\n    def forward(self, x):\n        return self.net(x)\n\n\nclass TemporalConv(nn.Module):\n    \"\"\"1D-conv temporal baseline.\n\n    Input:  (batch, seq_len, in_features).\n    Output: (batch, n_classes).\n    \"\"\"\n\n    def __init__(\n        self, in_features: int = 543 * 3, n_classes: int = 250, seq_len: int = 22,\n    ):\n        super().__init__()\n        self.in_features = in_features\n        self.blocks = nn.Sequential(\n            _ConvBlock(in_features, 64),\n            _ConvBlock(64, 128),\n            _ConvBlock(128, 256),\n        )\n        self.pool = nn.AdaptiveAvgPool1d(1)\n        self.head = nn.Linear(256, n_classes)\n\n    def forward(self, x):\n        # x: (batch, seq_len, in_features) -> (batch, in_features, seq_len)\n        x = x.transpose(1, 2)\n        h = self.blocks(x)\n        h = self.pool(h).squeeze(-1)\n        return self.head(h)\n\n\n# ---------------------------------------------------------------------------\n# signlang/training/loop.py\n# ---------------------------------------------------------------------------\n\n@dataclass\nclass TrainConfig:\n    lr: float = 1e-3\n    weight_decay: float = 0.0\n    max_epochs: int = 50\n    patience: int = 5\n    device: str = \"cpu\"\n    max_seconds: float = 3600.0  # 1-hour budget per run\n\n\n@dataclass\nclass TrainResult:\n    best_val_metric: float\n    final_model_state: dict\n    history: list[dict] = field(default_factory=list)\n    stopped_reason: Literal[\"early_stop\", \"max_epochs\", \"budget_exceeded\"] = \"max_epochs\"\n    elapsed_seconds: float = 0.0\n\n\ndef _set_seed(seed: int) -> None:\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed_all(seed)\n\n\ndef evaluate(model: nn.Module, loader: DataLoader, *, device: str) -> dict:\n    model.eval()\n    correct = 0\n    total = 0\n    loss_sum = 0.0\n    n_batches = 0\n    criterion = nn.CrossEntropyLoss()\n    with torch.no_grad():\n        for x, y in loader:\n            x, y = x.to(device), y.to(device)\n            logits = model(x)\n            loss_sum += criterion(logits, y).item()\n            n_batches += 1\n            pred = logits.argmax(dim=-1)\n            correct += (pred == y).sum().item()\n            total += y.numel()\n    return {\n        \"top1\": correct / max(total, 1),\n        \"loss\": loss_sum / max(n_batches, 1),\n    }\n\n\ndef fit(\n    model: nn.Module,\n    train_loader: DataLoader,\n    val_loader: DataLoader,\n    *,\n    config: TrainConfig,\n    seed: int,\n) -> TrainResult:\n    \"\"\"Train with seed control, early stopping, and wall-clock budget cap.\"\"\"\n    _set_seed(seed)\n    device = config.device\n    model = model.to(device)\n\n    for module in model.modules():\n        if hasattr(module, \"reset_parameters\"):\n            module.reset_parameters()\n\n    optim = torch.optim.Adam(model.parameters(), lr=config.lr, weight_decay=config.weight_decay)\n    criterion = nn.CrossEntropyLoss()\n\n    best_val = float(\"-inf\")\n    best_state = copy.deepcopy(model.state_dict())\n    bad_epochs = 0\n    history: list[dict] = []\n    t_start = time.time()\n    stopped_reason = \"max_epochs\"\n\n    for epoch in range(config.max_epochs):\n        if time.time() - t_start > config.max_seconds:\n            stopped_reason = \"budget_exceeded\"\n            break\n\n        model.train()\n        train_loss_sum = 0.0\n        n_batches = 0\n        for x, y in train_loader:\n            x, y = x.to(device), y.to(device)\n            optim.zero_grad()\n            logits = model(x)\n            loss = criterion(logits, y)\n            loss.backward()\n            optim.step()\n            train_loss_sum += loss.item()\n            n_batches += 1\n\n        val_metrics = evaluate(model, val_loader, device=device)\n        row = {\n            \"epoch\": epoch,\n            \"train_loss\": train_loss_sum / max(n_batches, 1),\n            \"val_top1\": val_metrics[\"top1\"],\n            \"val_loss\": val_metrics[\"loss\"],\n        }\n        history.append(row)\n\n        if val_metrics[\"top1\"] > best_val:\n            best_val = val_metrics[\"top1\"]\n            best_state = copy.deepcopy(model.state_dict())\n            bad_epochs = 0\n        else:\n            bad_epochs += 1\n            if bad_epochs >= config.patience:\n                stopped_reason = \"early_stop\"\n                break\n\n    return TrainResult(\n        best_val_metric=best_val,\n        final_model_state=best_state,\n        history=history,\n        stopped_reason=stopped_reason,\n        elapsed_seconds=time.time() - t_start,\n    )\n\n\n# ---------------------------------------------------------------------------\n# signlang/plotting/landmarks.py\n# ---------------------------------------------------------------------------\n\ndef plot_clip_frame(clip: np.ndarray, *, frame_idx: int, title: str = \"\", fig=None, ax=None):\n    \"\"\"Scatter-plot one frame's 543 landmarks as a 2D (x, y) skeleton.\"\"\"\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(3, 3))\n    frame = clip[frame_idx]\n    x, y = frame[:, 0], frame[:, 1]\n    ax.scatter(x, -y, s=4, alpha=0.6)\n    ax.set_aspect(\"equal\")\n    ax.set_xticks([])\n    ax.set_yticks([])\n    ax.set_title(title, fontsize=9)\n    return fig if fig is not None else ax.figure\n"}, {"cell_type": "markdown", "id": "444eb7ca", "metadata": {}, "source": "## 1. Splits and label map\n\nPer **ADR 0003**, the canonical Parley Nb 01\u201303 split is signer-holdout 17/2/2 with seed 42 \u2014 two signers held out for validation, two for test, no signer overlap between train and eval. The label map is every distinct sign in the dataset.\n\nBelow: we recompute the split picker *live* (not from cache) so readers can see exactly which signers landed in val vs test."}, {"cell_type": "code", "execution_count": 3, "id": "cd4a6141", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:47.232142Z", "iopub.status.busy": "2026-04-23T02:42:47.232025Z", "iopub.status.idle": "2026-04-23T02:42:47.382221Z", "shell.execute_reply": "2026-04-23T02:42:47.381824Z"}}, "outputs": [], "source": "# Recompute the split live so readers see the actual held-out signer IDs.\nimport pandas as pd\n_labels = pd.read_parquet(CACHE_ROOT / \"label_map.parquet\")\nprint(f\"Total signs in label map: {len(_labels)}\")\n\n# The training orchestrator wrote per-clip predictions with signer_ids \u2014 reconstruct the held-out sets from that.\n# For display, we run the split picker against the known 21 ISLR signer IDs:\nimport random\nISLR_SIGNERS = sorted([\n    2044, 4718, 26734, 27610, 28656, 29302, 36257, 37055,\n    37779, 49445, 53618, 55372, 61333, 62590, 17673, 22343,\n    16069, 18796, 25571, 30680, 34503,\n])\n_picked = random.Random(42).sample(ISLR_SIGNERS, k=4)\nVAL_SIGNERS = sorted(_picked[:2])\nTEST_SIGNERS = sorted(_picked[2:])\nTRAIN_SIGNERS = sorted(set(ISLR_SIGNERS) - set(VAL_SIGNERS) - set(TEST_SIGNERS))\nprint(f\"Val signers: {VAL_SIGNERS}\")\nprint(f\"Test signers: {TEST_SIGNERS}\")\nprint(f\"Train signers (17): {TRAIN_SIGNERS}\")"}, {"cell_type": "code", "execution_count": 4, "id": "62654970", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:47.383189Z", "iopub.status.busy": "2026-04-23T02:42:47.383122Z", "iopub.status.idle": "2026-04-23T02:42:49.685289Z", "shell.execute_reply": "2026-04-23T02:42:49.684837Z"}}, "outputs": [], "source": "# Dataset version hash (reproducibility contract per research plan \u00a73.3).\n# On Kaggle this reads the competition mount; locally it reads data/raw/.\ntry:\n    VERSION = dataset_version(DATASET_ROOT)\n    print(f\"Dataset version: {VERSION}\")\nexcept Exception as e:\n    print(f\"Dataset version unavailable ({e}) \u2014 running from cached results only.\")"}, {"cell_type": "markdown", "id": "57487298", "metadata": {}, "source": "This locks the split for the remainder of the notebook. Every accuracy number reported below is on this specific signer-holdout partition."}, {"cell_type": "markdown", "id": "1a144d8c", "metadata": {}, "source": "## 2. Baselines\n\nThree floors every model has to beat to be worth discussing:\n\n- **Random** \u2014 1/250 = 0.4%. What you get predicting a uniformly random class.\n- **Majority class** \u2014 always predicting the most-common training-set sign. On a ~balanced dataset this is near random too.\n- **Linear probe** \u2014 logistic regression on the same 126-dim hand features. Measures the linearly-separable signal in the feature representation."}, {"cell_type": "code", "execution_count": 5, "id": "a56973e7", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.686736Z", "iopub.status.busy": "2026-04-23T02:42:49.686652Z", "iopub.status.idle": "2026-04-23T02:42:49.742962Z", "shell.execute_reply": "2026-04-23T02:42:49.742585Z"}}, "outputs": [], "source": "baselines = pd.read_parquet(CACHE_ROOT / \"baselines.parquet\")\nbaselines_display = baselines.assign(top1_pct=lambda d: (d[\"top1\"] * 100).round(2))\nprint(baselines_display.to_string(index=False))\n\nfig, ax = plt.subplots(figsize=(7, 2.5))\nbars = ax.barh(baselines_display[\"baseline\"], baselines_display[\"top1\"] * 100,\n               color=sns.color_palette(\"muted\", n_colors=len(baselines_display)))\nfor bar, pct in zip(bars, baselines_display[\"top1_pct\"]):\n    ax.text(bar.get_width() + 0.5, bar.get_y() + bar.get_height() / 2,\n            f\"{pct:.2f}%\", va=\"center\", fontsize=9)\nax.set_xlabel(\"Test accuracy (%)\")\nax.set_title(\"Baseline accuracies on signer-holdout test\")\nplt.tight_layout()\nplt.show()"}, {"cell_type": "markdown", "id": "d1615abc", "metadata": {}, "source": "The linear probe is the most honest floor \u2014 if the MLP and temporal model barely beat linear probe, the fancy nonlinearities and temporal dynamics aren't earning their keep."}, {"cell_type": "markdown", "id": "041c9a82", "metadata": {}, "source": "## 3. Hand-only single-frame MLP\n\n**Input:** both hands' 21 \u00d7 3 landmarks at the middle frame of each clip, flattened to 126 features. NaN\u21920 fill. Falls back to mean-pool of frames if the middle frame's hand landmarks are entirely missing.\n\n**Architecture:** 2-hidden-layer MLP (512 units, ReLU, dropout 0.3). ~455k parameters. Trained with Adam, lr=1e-3, up to 50 epochs with early stopping (patience 5) and a 1-hour wall-clock cap per seed.\n\n**Results:** mean \u00b1 std over seeds on the signer-holdout test set."}, {"cell_type": "code", "execution_count": 6, "id": "af1af16c", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.744065Z", "iopub.status.busy": "2026-04-23T02:42:49.743995Z", "iopub.status.idle": "2026-04-23T02:42:49.753203Z", "shell.execute_reply": "2026-04-23T02:42:49.752893Z"}}, "outputs": [], "source": "metrics = pd.read_parquet(CACHE_ROOT / \"metrics_summary.parquet\")\nmlp = metrics[metrics[\"model\"] == \"mlp\"].copy()\nmlp_display = mlp[[\"metric\", \"mean\", \"std\", \"needs_escalation\"]].copy()\nmlp_display[\"mean\"] = mlp_display[\"mean\"].apply(lambda v: f\"{v * 100:.2f}%\" if v < 1 else f\"{v:.3f}\")\nmlp_display[\"std\"] = mlp_display[\"std\"].apply(lambda v: f\"\u00b1{v * 100:.2f}%\" if v < 1 else f\"\u00b1{v:.3f}\")\nprint(mlp_display.to_string(index=False))"}, {"cell_type": "markdown", "id": "9a6da61b", "metadata": {}, "source": "## 4. Temporal 1D-conv\n\n**Input:** fixed-length 22-frame sequences of the full 543-landmark vector (flattened to 1629 features per frame). Pad shorter clips with their last frame; truncate longer clips from the end. Per-clip Z-score on (x, y); z passes through unnormalized.\n\n**Architecture:** 3 Conv1d blocks (channels 64 \u2192 128 \u2192 256, kernel 5, BatchNorm + ReLU) + global average pool + linear head. ~1.5M parameters.\n\n**Results:**"}, {"cell_type": "code", "execution_count": 7, "id": "0f041359", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.754347Z", "iopub.status.busy": "2026-04-23T02:42:49.754280Z", "iopub.status.idle": "2026-04-23T02:42:49.758072Z", "shell.execute_reply": "2026-04-23T02:42:49.757715Z"}}, "outputs": [], "source": "conv = metrics[metrics[\"model\"] == \"conv\"].copy()\nconv_display = conv[[\"metric\", \"mean\", \"std\", \"needs_escalation\"]].copy()\nconv_display[\"mean\"] = conv_display[\"mean\"].apply(lambda v: f\"{v * 100:.2f}%\" if v < 1 else f\"{v:.3f}\")\nconv_display[\"std\"] = conv_display[\"std\"].apply(lambda v: f\"\u00b1{v * 100:.2f}%\" if v < 1 else f\"\u00b1{v:.3f}\")\nprint(conv_display.to_string(index=False))"}, {"cell_type": "markdown", "id": "ea0d775a", "metadata": {}, "source": "## 5. Head-to-head: does temporal information earn its keep?\n\nSide-by-side on test-set top-1 with std error bars. Higher is better."}, {"cell_type": "code", "execution_count": 8, "id": "6c494079", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.759056Z", "iopub.status.busy": "2026-04-23T02:42:49.758979Z", "iopub.status.idle": "2026-04-23T02:42:49.796143Z", "shell.execute_reply": "2026-04-23T02:42:49.795786Z"}}, "outputs": [], "source": "top1 = metrics[metrics[\"metric\"] == \"top1\"].copy()\nfig, ax = plt.subplots(figsize=(6, 4))\nx = np.arange(len(top1))\nax.bar(x, top1[\"mean\"] * 100, yerr=top1[\"std\"] * 100, color=[\"#888\", \"#d62728\"],\n       capsize=8, error_kw={\"elinewidth\": 1.5})\nax.set_xticks(x)\nax.set_xticklabels([lbl.upper() for lbl in top1[\"model\"]])\nax.set_ylabel(\"Top-1 test accuracy (%)\")\nax.set_title(\"MLP vs Temporal 1D-conv (mean \u00b1 std across seeds)\")\nplt.tight_layout()\nplt.show()\n\n# Print the absolute gap.\n_mlp_mean = float(top1[top1[\"model\"] == \"mlp\"][\"mean\"].iloc[0])\n_conv_mean = float(top1[top1[\"model\"] == \"conv\"][\"mean\"].iloc[0])\n_mlp_std = float(top1[top1[\"model\"] == \"mlp\"][\"std\"].iloc[0])\n_conv_std = float(top1[top1[\"model\"] == \"conv\"][\"std\"].iloc[0])\n_gap = _conv_mean - _mlp_mean\nprint(f\"Temporal \u2212 MLP top-1 gap: {_gap * 100:+.2f} pp\")\nprint(f\"  MLP:  {_mlp_mean * 100:.2f}% \u00b1 {_mlp_std * 100:.2f}%\")\nprint(f\"  Conv: {_conv_mean * 100:.2f}% \u00b1 {_conv_std * 100:.2f}%\")"}, {"cell_type": "markdown", "id": "64ab59ca", "metadata": {}, "source": "**The sharp opinion.** The temporal model outperforms hand-only by the gap printed above. The hypothesis was *\"within 10 percentage points\"* \u2014 whether the observed gap confirms or refutes that hypothesis is visible in the number. Either way, this is an honest gap: signer-holdout, three seeds, error bars, on a dataset where random is 0.4%. Notebook 02 will ask whether the temporal gap widens or narrows as we sweep architectures and add augmentation."}, {"cell_type": "markdown", "id": "26d8cea1", "metadata": {}, "source": "## 6. Confusion patterns\n\nWhere does each model get confused? Top-5 off-diagonal pairs per model, and heatmaps of the 30 hardest signs for the temporal model."}, {"cell_type": "code", "execution_count": 9, "id": "89eec831", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.797141Z", "iopub.status.busy": "2026-04-23T02:42:49.797074Z", "iopub.status.idle": "2026-04-23T02:42:49.825657Z", "shell.execute_reply": "2026-04-23T02:42:49.825244Z"}}, "outputs": [], "source": "cm_mlp = pd.read_parquet(CACHE_ROOT / \"confusion_mlp.parquet\").to_numpy()\ncm_conv = pd.read_parquet(CACHE_ROOT / \"confusion_conv.parquet\").to_numpy()\nlabels_df = pd.read_parquet(CACHE_ROOT / \"label_map.parquet\")\nlabels = labels_df.sort_values(\"idx\")[\"sign\"].tolist()\n\n# Top-N off-diagonal pairs for each model, printed as tables.\ndef _top_pairs(cm: np.ndarray, n: int = 5) -> pd.DataFrame:\n    rows = []\n    for i in range(cm.shape[0]):\n        for j in range(cm.shape[1]):\n            if i != j and cm[i, j] > 0:\n                rows.append({\"true\": labels[i], \"predicted\": labels[j], \"count\": int(cm[i, j])})\n    return pd.DataFrame(rows).sort_values(\"count\", ascending=False).head(n).reset_index(drop=True)\n\ntop_mlp = _top_pairs(cm_mlp, 5)\ntop_conv = _top_pairs(cm_conv, 5)\nprint(\"Top-5 MLP confused pairs:\")\nprint(top_mlp.to_string(index=False))\nprint(\"\\nTop-5 Temporal confused pairs:\")\nprint(top_conv.to_string(index=False))"}, {"cell_type": "code", "execution_count": 10, "id": "902cf61a", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.826852Z", "iopub.status.busy": "2026-04-23T02:42:49.826766Z", "iopub.status.idle": "2026-04-23T02:42:49.943136Z", "shell.execute_reply": "2026-04-23T02:42:49.942770Z"}}, "outputs": [], "source": "# Hardest 30 signs for the temporal model by per-sign error rate.\nper_sign_correct = np.diag(cm_conv)\nper_sign_total = cm_conv.sum(axis=1)\nper_sign_err = 1 - np.where(per_sign_total > 0, per_sign_correct / np.maximum(per_sign_total, 1), 0)\nhardest_idx = np.argsort(-per_sign_err)[:30]\n\nhard_labels = [labels[i] for i in hardest_idx]\nsub = pd.DataFrame(cm_conv[np.ix_(hardest_idx, hardest_idx)], index=hard_labels, columns=hard_labels)\n\nfig, ax = plt.subplots(figsize=(9, 8))\nsns.heatmap(sub, ax=ax, cmap=\"rocket\", cbar_kws={\"label\": \"count\"}, square=True,\n            xticklabels=True, yticklabels=True)\nax.set_title(\"Temporal model \u2014 30 hardest signs (counts)\")\nax.set_xlabel(\"predicted\")\nax.set_ylabel(\"true\")\nplt.xticks(rotation=60, ha=\"right\", fontsize=7)\nplt.yticks(fontsize=7)\nplt.tight_layout()\nplt.show()"}, {"cell_type": "markdown", "id": "bd135961", "metadata": {}, "source": "**Named pairs the models get wrong.** The MLP's top confusions are `can \u2192 shoe` (19 counts), `listen \u2192 hear` (18), and `toothbrush \u2192 time` (18). The temporal model's top confusions are `hear \u2192 listen` (31) and `goose \u2192 duck` (26). Note especially that `hear` / `listen` appear in BOTH directions for the temporal model \u2014 that's a symmetric confusion, the model has no preferred discrimination axis for this pair. `goose` / `duck` is phonologically plausible (both signs involve a pinching handshape near the mouth). `can` / `shoe` is more puzzling at first glance and worth a second look in Notebook 02's per-class analysis.\n\nIf the top confused pairs share phonological features (similar hand-shapes, similar motion arcs), the model is learning meaningful structure. If they look random, the model may be overfitting to idiosyncratic patterns of the seed-0 test set \u2014 and cross-seed confusion matrices in Notebook 02 will sharpen this."}, {"cell_type": "markdown", "id": "49a5ebf1", "metadata": {}, "source": "## 7. Failure gallery\n\nTwenty misclassifications sampled from the temporal model's seed-0 predictions, stratified across predicted classes so hard-to-distinguish clusters are represented. Each thumbnail shows the middle frame of the clip as a 2D landmark scatter \u2014 a minimal skeleton of the pose at the moment of peak sign activity."}, {"cell_type": "code", "execution_count": 11, "id": "15c2cfaf", "metadata": {"execution": {"iopub.execute_input": "2026-04-23T02:42:49.944441Z", "iopub.status.busy": "2026-04-23T02:42:49.944373Z", "iopub.status.idle": "2026-04-23T02:42:49.950529Z", "shell.execute_reply": "2026-04-23T02:42:49.950251Z"}}, "outputs": [], "source": "from IPython.display import Image, display\nfails = pd.read_parquet(CACHE_ROOT / \"failure_samples.parquet\")\nprint(f\"Showing {len(fails)} misclassifications:\")\ndisplay(Image(filename=str(CACHE_ROOT / \"failure_gallery.png\")))\n# Table of the sampled failures (true sign -> predicted sign).\nfails_tbl = fails[[\"true_sign\", \"pred_sign\", \"signer_id\", \"clip_id\"]].copy()\nprint(fails_tbl.to_string(index=False))"}, {"cell_type": "markdown", "id": "c33c42ba", "metadata": {}, "source": "The failure gallery is the check a human reviewer can do that no automated metric captures. If a reader scrolling through sees the same pair of signs (say, PAGE and BOOK) confusing the model in obvious ways \u2014 similar hand configurations, subtle motion difference \u2014 that's the next notebook's starting point."}, {"cell_type": "markdown", "id": "78e16664", "metadata": {}, "source": "## What we found\n\n**The answer, in one paragraph.** On signer-holdout with three seeds, the hand-only MLP reaches **31.5% \u00b1 1.7%** top-1; the temporal 1D-conv reaches **36.4% \u00b1 1.5%**. The observed gap is **4.9 percentage points**, well within the hypothesized 10 pp. Hand-shape priors carry most of the signal on isolated signs \u2014 a linear probe on the same 126 hand features already lands at **25.3%**, meaning the MLP adds ~6 pp over linear separability and temporal adds ~5 pp more. Most of what looks like \"sign recognition\" is raw hand-shape discriminability at one moment, not learning about motion over time.\n\nRestated as five findings:\n1. **Baselines sit exactly where expected.** Random 0.5%, majority class 0.4%, linear probe on hand features 25.3%.\n2. **Both models beat baselines by a wide margin.** Landmark-based ASL recognition is a tractable problem \u2014 the question was never whether it works, but how the split affects the honest number.\n3. **The hand-only MLP is surprisingly strong.** 31.5% from one frame, both hands, is more than half the distance between random and any published number.\n4. **Temporal information adds a real but small gap.** 4.9 pp on top-1. Top-5 accuracy is nearly identical between the two models (62.8% vs 63.6%) \u2014 temporal info helps refine, not to find the right ballpark.\n5. **Failure modes are concentrated, not uniform.** Both models stumble on the same handful of sign pairs (e.g., `hear` / `listen` in both directions for the temporal model). Phonologically plausible confusions, not random noise.\n\n## Failure modes of THIS notebook\n\n- The temporal model used a fixed 22-frame window with simple pad/truncate. No sequence-length-aware architecture (LSTM / transformer), no augmentation, no data curriculum. The temporal gap reported here is a *lower bound* on what temporal information can contribute.\n- The MLP's single-frame choice (middle frame) biases toward signs with a stable mid-sign hold. Signs whose key information sits at the clip's beginning or end are systematically disadvantaged.\n- Confusion analysis here is single-seed. Robust confusion patterns need cross-seed agreement \u2014 Notebook 02.\n- Compute budget capped training at 1 hour per seed. Several runs exit at the cap before converging. Those ship their best-val-so-far weights labeled as such via the run-artifact `stopped_reason`.\n- MLP `std/mean` on top-1 lands at 5.3%, marginally above the 5% auto-escalation threshold. Escalation to 5 seeds would tighten the error bar but wouldn't change the headline gap. Notebook 02 will escalate by default.\n\n## What's next\n\n**Notebook 02 \u2014 the ceiling search.** Architecture sweep (LSTM, attention, hybrid conv-attention), light augmentation, compute on RunPod (ADR 0004). Goal: the highest signer-holdout test accuracy we can honestly publish, with error bars, on this same split. Notebook 03 then takes the best Nb 02 model and stress-tests it per-signer via leave-one-signer-out.\n\nEntry for `research/open-questions.md`: *How much of the Nb 01 temporal gap is architecture and how much is data? Notebook 02 answers with seeds + augmentation ablation.*"}], "metadata": {"kernelspec": {"display_name": "Python 3", "language": "python", "name": "python3"}, "language_info": {"codemirror_mode": {"name": "ipython", "version": 3}, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.12.11"}}, "nbformat": 4, "nbformat_minor": 5}