{"cells": [{"cell_type": "markdown", "id": "2f391960", "metadata": {}, "source": "# What does the Google ASL Signs data actually look like?\n\n*A plain-eyes walk through 94,000 clips, 250 signs, 21 signers, and 543 landmarks per frame \u2014 with the questions a model is about to try to answer written down before we open any of it.*\n\n**Question.** What is the shape, structure, missing-data profile, signer coverage, and data-quality story of Google's Isolated Sign Language Recognition landmark dataset \u2014 enough that every downstream Parley notebook can inherit a coherent picture of what it is modeling?\n\n**Hypothesis.** The dataset has enough systematic structure (non-uniform per-signer coverage, non-random missing landmarks, meaningful variance in clip length) that the naive Kaggle-default random split hides signer leakage and inflates accuracy; a signer-holdout split is the honest default.\n\n**Dataset.** Google Isolated Sign Language Recognition (Kaggle competition `asl-signs`). 250 ASL signs, 21 signers, pre-extracted MediaPipe landmarks \u2014 no raw video. Landmark parquet, ~37 GB extracted. **License:** Kaggle competition terms; research use permitted, bulk redistribution is not.\n\n**Prior work.** Winning Kaggle entries (~0.83 public LB) use temporal CNNs or transformers on landmark sequences (e.g., Kazuki Onodera, 2023). None of the published solutions we could find report signer-holdout numbers alongside their competition score.\n\n**How to reproduce.** This notebook runs standalone on Kaggle \u2014 the `asl-signs` competition data and a pre-computed aggregation cache (`parley-nb00-caches`) are both attached. Clicking *Run All* regenerates every figure in under a minute. For the full source (modular `signlang/` package, tests, and cold-rebuild from the raw 94k parquets) see the Parley repo.\n"}, {"cell_type": "code", "execution_count": 1, "id": "380ca1d9", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:47:59.468097Z", "iopub.status.busy": "2026-04-22T15:47:59.467879Z", "iopub.status.idle": "2026-04-22T15:48:00.747787Z", "shell.execute_reply": "2026-04-22T15:48:00.747405Z"}}, "outputs": [], "source": "%matplotlib inline\nfrom __future__ import annotations\n\nfrom dataclasses import dataclass\nfrom pathlib import Path\nfrom typing import Callable\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\n\n# ---------------------------------------------------------------------------\n# Path detection \u2014 Kaggle runtime vs local repo.\n# ---------------------------------------------------------------------------\n_KAGGLE_INPUT = Path(\"/kaggle/input\")\nIN_KAGGLE = _KAGGLE_INPUT.exists()\n\nif IN_KAGGLE:\n    # Competition data mounted at /kaggle/input/asl-signs/; caches at\n    # /kaggle/input/parley-nb00-caches/.\n    DATASET_ROOT = _KAGGLE_INPUT / \"asl-signs\"\n    CACHE_ROOT = _KAGGLE_INPUT / \"parley-nb00-caches\"\nelse:\n    try:\n        NB_DIR = Path(__file__).resolve().parent  # type: ignore[name-defined]\n    except NameError:\n        NB_DIR = Path.cwd() if Path.cwd().name == \"notebooks\" else Path.cwd() / \"notebooks\"\n    REPO_ROOT = NB_DIR.parent\n    DATASET_ROOT = REPO_ROOT / \"data\" / \"raw\" / \"google-asl-signs\"\n    CACHE_ROOT = REPO_ROOT / \"data\" / \"processed\" / \"notebook-00\"\n\n# ---------------------------------------------------------------------------\n# Analysis parameters (hardcoded; config file is not available on Kaggle).\n# ---------------------------------------------------------------------------\nSEED = 42\nMANIFEST = \"train.csv\"\nMIN_FRAMES_PER_CLIP = 4\nHEATMAP_CLIPS_PER_SIGN = 20\n\nsns.set_theme(style=\"whitegrid\", context=\"notebook\")\nnp.random.seed(SEED)\n\nprint(\"Running on:\", \"Kaggle\" if IN_KAGGLE else \"local repo\")\n"}, {"cell_type": "code", "execution_count": 2, "id": "a9928e34", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:00.748972Z", "iopub.status.busy": "2026-04-22T15:48:00.748866Z", "iopub.status.idle": "2026-04-22T15:48:00.756891Z", "shell.execute_reply": "2026-04-22T15:48:00.756643Z"}}, "outputs": [], "source": "# =========================================================================\n# Inlined from Parley's `signlang/` package so this notebook runs standalone\n# on Kaggle. No logic differences from the modular source \u2014 the repo keeps\n# it split across signlang/data/ and signlang/analysis/ with pytest coverage.\n# =========================================================================\n\n\n@dataclass(frozen=True)\nclass ClipMetadata:\n    \"\"\"Lightweight metadata for one ISLR clip.\"\"\"\n    clip_id: str\n    sign: str\n    signer_id: int\n    frame_count: int\n    missing_rate: float\n\n\nLANDMARK_TYPE_COUNTS = {\"face\": 468, \"left_hand\": 21, \"pose\": 33, \"right_hand\": 21}\nLANDMARK_ORDER = (\"face\", \"left_hand\", \"pose\", \"right_hand\")\nN_LANDMARKS = sum(LANDMARK_TYPE_COUNTS.values())  # 543\n\n\ndef load_clip(path) -> np.ndarray:\n    \"\"\"Load one ISLR parquet as a (n_frames, 543, 3) float32 array.\"\"\"\n    df = pd.read_parquet(path)\n    frames = np.sort(df[\"frame\"].unique())\n    n_frames = len(frames)\n    frame_to_row = {f: i for i, f in enumerate(frames)}\n    out = np.full((n_frames, N_LANDMARKS, 3), np.nan, dtype=np.float32)\n    offsets = {}\n    cursor = 0\n    for lm_type in LANDMARK_ORDER:\n        offsets[lm_type] = cursor\n        cursor += LANDMARK_TYPE_COUNTS[lm_type]\n    for lm_type in LANDMARK_TYPE_COUNTS:\n        sub = df[df[\"type\"] == lm_type]\n        if sub.empty:\n            continue\n        rows = np.array([frame_to_row[f] for f in sub[\"frame\"].to_numpy()])\n        cols = offsets[lm_type] + sub[\"landmark_index\"].to_numpy(dtype=np.int64)\n        out[rows, cols, 0] = sub[\"x\"].to_numpy(dtype=np.float32)\n        out[rows, cols, 1] = sub[\"y\"].to_numpy(dtype=np.float32)\n        out[rows, cols, 2] = sub[\"z\"].to_numpy(dtype=np.float32)\n    return out\n\n\ndef missing_mask(clip: np.ndarray) -> np.ndarray:\n    \"\"\"True = missing detection (NaN in any coord, OR all-zero row).\"\"\"\n    if clip.ndim != 3 or clip.shape[1] != N_LANDMARKS or clip.shape[2] != 3:\n        raise ValueError(f\"expected (n_frames, {N_LANDMARKS}, 3), got {clip.shape}\")\n    nan_missing = np.isnan(clip).any(axis=-1)\n    zero_missing = (clip == 0.0).all(axis=-1)\n    return nan_missing | zero_missing\n\n\ndef list_clips(root, *, manifest_name=\"train.csv\") -> list:\n    \"\"\"Return a ClipMetadata per manifest row (opens each parquet \u2014 O(N-clips)).\"\"\"\n    root = Path(root)\n    manifest = pd.read_csv(root / manifest_name)\n    out = []\n    for _, row in manifest.iterrows():\n        clip = load_clip(root / row[\"path\"])\n        n_frames = clip.shape[0]\n        missing = float(missing_mask(clip).mean()) if n_frames > 0 else 0.0\n        out.append(ClipMetadata(\n            clip_id=str(row[\"sequence_id\"]),\n            sign=str(row[\"sign\"]),\n            signer_id=int(row[\"participant_id\"]),\n            frame_count=int(n_frames),\n            missing_rate=float(missing),\n        ))\n    return out\n\n\ndef full_dataset_stats(root, *, manifest_name=\"train.csv\") -> pd.DataFrame:\n    \"\"\"Long-format DataFrame of dataset-wide summary metrics (metric, value).\"\"\"\n    clips = list_clips(root, manifest_name=manifest_name)\n    if not clips:\n        raise ValueError(f\"no clips found under {root}\")\n    frame_counts = np.asarray([c.frame_count for c in clips], dtype=np.int64)\n    missing_rates = np.asarray([c.missing_rate for c in clips], dtype=np.float64)\n    signs = {c.sign for c in clips}\n    signers = {c.signer_id for c in clips}\n    rows = [\n        (\"total_clips\", float(len(clips))),\n        (\"total_signs\", float(len(signs))),\n        (\"total_signers\", float(len(signers))),\n        (\"mean_frames_per_clip\", float(frame_counts.mean())),\n        (\"median_frames_per_clip\", float(np.median(frame_counts))),\n        (\"min_frames_per_clip\", float(frame_counts.min())),\n        (\"max_frames_per_clip\", float(frame_counts.max())),\n        (\"overall_missing_rate\", float(missing_rates.mean())),\n    ]\n    return pd.DataFrame(rows, columns=[\"metric\", \"value\"])\n\n\ndef load_or_compute(cache_path, compute_fn, *, force=False) -> pd.DataFrame:\n    \"\"\"Read parquet cache or compute+write on miss. Core Kaggle-fast-path pattern.\"\"\"\n    cache_path = Path(cache_path)\n    if cache_path.exists() and not force:\n        return pd.read_parquet(cache_path)\n    cache_path.parent.mkdir(parents=True, exist_ok=True)\n    result = compute_fn()\n    result.to_parquet(cache_path)\n    return result\n\n\ndef clips_per_sign(clips) -> pd.Series:\n    signs = pd.Series([c.sign for c in clips], name=\"clips\")\n    out = signs.value_counts().sort_values(ascending=False)\n    out.name = \"clips\"\n    out.index.name = \"sign\"\n    return out\n\n\ndef clips_per_signer(clips) -> pd.Series:\n    ids = pd.Series([c.signer_id for c in clips], name=\"clips\")\n    out = ids.value_counts().sort_values(ascending=False)\n    out.name = \"clips\"\n    out.index.name = \"signer_id\"\n    return out\n\n\ndef frames_per_clip(clips) -> pd.Series:\n    return pd.Series([c.frame_count for c in clips], name=\"frame_count\")\n\n\ndef per_signer_frame_count_summary(clips) -> pd.DataFrame:\n    df = pd.DataFrame({\n        \"signer_id\": [c.signer_id for c in clips],\n        \"frame_count\": [c.frame_count for c in clips],\n    })\n    g = df.groupby(\"signer_id\")[\"frame_count\"]\n    return pd.DataFrame({\n        \"count\": g.count(), \"mean\": g.mean(), \"std\": g.std().fillna(0.0),\n        \"median\": g.median(), \"min\": g.min(), \"max\": g.max(),\n    })\n\n\ndef missing_rate_by_type(clip: np.ndarray) -> dict:\n    mask = missing_mask(clip)\n    out = {}\n    cursor = 0\n    for lm_type in LANDMARK_ORDER:\n        count = LANDMARK_TYPE_COUNTS[lm_type]\n        block = mask[:, cursor : cursor + count]\n        out[lm_type] = float(block.mean()) if block.size > 0 else 0.0\n        cursor += count\n    return out\n\n\ndef per_signer_missing_rate(clips) -> pd.Series:\n    df = pd.DataFrame({\n        \"signer_id\": [c.signer_id for c in clips],\n        \"missing_rate\": [c.missing_rate for c in clips],\n    })\n    return df.groupby(\"signer_id\")[\"missing_rate\"].mean().rename(\"missing_rate\")\n\n\ndef sign_landmark_type_missing_matrix(clips, *, load_clip_fn, clips_per_sign_cap=None) -> pd.DataFrame:\n    grouped = {}\n    for c in clips:\n        grouped.setdefault(c.sign, []).append(c)\n    rows = []\n    for sign, sign_clips in grouped.items():\n        if clips_per_sign_cap is not None:\n            sign_clips = sign_clips[:clips_per_sign_cap]\n        accum = {t: [] for t in LANDMARK_ORDER}\n        for c in sign_clips:\n            rel = f\"train_landmark_files/{c.signer_id}/{c.clip_id}.parquet\"\n            clip = load_clip_fn(rel)\n            rates = missing_rate_by_type(clip)\n            for t, r in rates.items():\n                accum[t].append(r)\n        row = {\"sign\": sign}\n        for t, rs in accum.items():\n            row[t] = float(np.mean(rs)) if rs else 0.0\n        rows.append(row)\n    return pd.DataFrame(rows).set_index(\"sign\")[list(LANDMARK_ORDER)]\n\n\ndef count_all_zero_frames(clip: np.ndarray) -> int:\n    if clip.ndim != 3 or clip.shape[2] != 3:\n        raise ValueError(f\"expected (n_frames, n_landmarks, 3), got {clip.shape}\")\n    m = np.nanmax(np.abs(clip), axis=(1, 2))\n    m = np.where(np.isnan(m), 1.0, m)\n    return int((m == 0.0).sum())\n\n\ndef count_duplicate_consecutive_frames(clip: np.ndarray) -> int:\n    if clip.ndim != 3:\n        raise ValueError(f\"expected 3-D, got {clip.shape}\")\n    if clip.shape[0] < 2:\n        return 0\n    a, b = clip[:-1], clip[1:]\n    eq_or_both_nan = (a == b) | (np.isnan(a) & np.isnan(b))\n    return int(eq_or_both_nan.all(axis=(1, 2)).sum())\n\n\ndef clip_quality_row(md, clip) -> dict:\n    return {\n        \"clip_id\": md.clip_id, \"sign\": md.sign, \"signer_id\": md.signer_id,\n        \"frame_count\": md.frame_count, \"missing_rate\": md.missing_rate,\n        \"all_zero_frames\": count_all_zero_frames(clip),\n        \"duplicate_consecutive_frames\": count_duplicate_consecutive_frames(clip),\n    }\n"}, {"cell_type": "markdown", "id": "0294be13", "metadata": {}, "source": "## 1. Shape counts\n\nHow much of what, belonging to whom. One cell answers the questions every reader will ask in the first ten seconds of seeing \"250 signs, 21 signers.\" We compute the shape once and cache it \u2014 the scan is ~6 min cold, instant warm.\n"}, {"cell_type": "code", "execution_count": 3, "id": "42cabc6f", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:00.757998Z", "iopub.status.busy": "2026-04-22T15:48:00.757943Z", "iopub.status.idle": "2026-04-22T15:48:00.832587Z", "shell.execute_reply": "2026-04-22T15:48:00.832195Z"}}, "outputs": [], "source": "stats_cache = CACHE_ROOT / \"shape_stats.parquet\"\n\n\ndef _compute_shape_stats() -> pd.DataFrame:\n    return full_dataset_stats(DATASET_ROOT, manifest_name=MANIFEST)\n\n\nshape_stats = load_or_compute(stats_cache, _compute_shape_stats)\nshape_stats\n"}, {"cell_type": "markdown", "id": "c48f9c44", "metadata": {}, "source": "*Read this left to right:* 250 signs across 21 signers, with roughly 94k clips and a mean clip length near 36 frames (\u22481.2 s at 30 fps). The max clip length is long enough that fixed-length padding will hurt; we'll come back to this in \u00a74.\n"}, {"cell_type": "code", "execution_count": 4, "id": "a7e6abab", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:00.833533Z", "iopub.status.busy": "2026-04-22T15:48:00.833470Z", "iopub.status.idle": "2026-04-22T15:48:01.057416Z", "shell.execute_reply": "2026-04-22T15:48:01.057038Z"}}, "outputs": [], "source": "clips_cache = CACHE_ROOT / \"clips_list.parquet\"\n\n\ndef _compute_clips_df() -> pd.DataFrame:\n    clips = list_clips(DATASET_ROOT, manifest_name=MANIFEST)\n    return pd.DataFrame(\n        [\n            {\n                \"clip_id\": c.clip_id,\n                \"sign\": c.sign,\n                \"signer_id\": c.signer_id,\n                \"frame_count\": c.frame_count,\n                \"missing_rate\": c.missing_rate,\n            }\n            for c in clips\n        ]\n    )\n\n\nclips_df = load_or_compute(clips_cache, _compute_clips_df)\n\n\ndef _clips_from_df(df: pd.DataFrame) -> list[ClipMetadata]:\n    return [ClipMetadata(**row) for row in df.to_dict(orient=\"records\")]\n\n\nCLIPS = _clips_from_df(clips_df)\nlen(CLIPS)\n"}, {"cell_type": "markdown", "id": "aa283e35", "metadata": {}, "source": "## 2. What a single frame contains\n\nMediaPipe emits 543 landmarks per frame for this dataset: a face mesh (468), a pose skeleton (33), a left hand (21), and a right hand (21). Every downstream model slices this 543-vector one way or another; this section names the slices so later cells don't have to.\n"}, {"cell_type": "code", "execution_count": 5, "id": "08204a65", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.058789Z", "iopub.status.busy": "2026-04-22T15:48:01.058716Z", "iopub.status.idle": "2026-04-22T15:48:01.100508Z", "shell.execute_reply": "2026-04-22T15:48:01.100111Z"}}, "outputs": [], "source": "fig, ax = plt.subplots(figsize=(7, 3.5))\ncounts = [LANDMARK_TYPE_COUNTS[t] for t in LANDMARK_ORDER]\ncolors = sns.color_palette(\"deep\", n_colors=len(LANDMARK_ORDER))\nbars = ax.barh(LANDMARK_ORDER, counts, color=colors)\nfor bar, count in zip(bars, counts):\n    ax.text(count + 5, bar.get_y() + bar.get_height() / 2, str(count),\n            va=\"center\", fontsize=10)\nax.set_xlabel(\"Landmarks per frame\")\nax.set_title(f\"543 landmarks per frame, by MediaPipe type\")\nax.set_xlim(0, max(counts) * 1.15)\nplt.tight_layout()\nplt.show()\n"}, {"cell_type": "code", "execution_count": 6, "id": "586eab59", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.101447Z", "iopub.status.busy": "2026-04-22T15:48:01.101383Z", "iopub.status.idle": "2026-04-22T15:48:01.104688Z", "shell.execute_reply": "2026-04-22T15:48:01.104381Z"}}, "outputs": [], "source": "# Our cached clips table: one row per clip with per-clip metadata.\n# frame_count confirms the dataset really has 543 landmarks per frame once\n# loaded (shape[1] == N_LANDMARKS) \u2014 the cache was produced by load_clip(),\n# which asserts exactly that shape on every parquet it reads.\nassert N_LANDMARKS == 543\nclips_df.head()\n"}, {"cell_type": "markdown", "id": "9c5e2b27", "metadata": {}, "source": "Missing-detection convention (from `signlang.data.missing_mask`): a landmark is treated as missing if **any** of (x, y, z) is NaN, or if **all** of (x, y, z) are exactly 0.0. The zero-fill case matters for pose rows in some clips; disambiguating signal-at-origin from zero-fill is impossible in the raw data, so we conservatively call both missing.\n"}, {"cell_type": "markdown", "id": "d30580b6", "metadata": {}, "source": "## 3. Where the data goes missing\n\nIf face landmarks are missing for signs where the hand passes over the face, or if one hand is reliably absent for bilateral signs, that's not random noise \u2014 that's a model-relevant signal about detection. We compute a (sign \u00d7 landmark-type) missing-rate matrix and plot the extremes.\n\nWe cap at `config.heatmap_clips_per_sign` clips per sign; with 250 signs and the default 20 clips each, that's 5000 parquet reads, ~90 s cold.\n"}, {"cell_type": "code", "execution_count": 7, "id": "8ba6ff51", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.105695Z", "iopub.status.busy": "2026-04-22T15:48:01.105636Z", "iopub.status.idle": "2026-04-22T15:48:01.112619Z", "shell.execute_reply": "2026-04-22T15:48:01.112253Z"}}, "outputs": [], "source": "heatmap_cache = CACHE_ROOT / \"sign_type_missing.parquet\"\n\n\ndef _compute_missing_matrix() -> pd.DataFrame:\n    return sign_landmark_type_missing_matrix(\n        CLIPS,\n        load_clip_fn=lambda rel: load_clip(DATASET_ROOT / rel),\n        clips_per_sign_cap=HEATMAP_CLIPS_PER_SIGN,\n    )\n\n\nmissing_matrix = load_or_compute(heatmap_cache, _compute_missing_matrix)\nmissing_matrix.describe()\n"}, {"cell_type": "code", "execution_count": 8, "id": "89455e9d", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.113455Z", "iopub.status.busy": "2026-04-22T15:48:01.113391Z", "iopub.status.idle": "2026-04-22T15:48:01.188002Z", "shell.execute_reply": "2026-04-22T15:48:01.187613Z"}}, "outputs": [], "source": "# Show the 30 signs with the widest spread across landmark types \u2014 the ones\n# where missingness is structural, not uniform.\nspread = missing_matrix.max(axis=1) - missing_matrix.min(axis=1)\ntop_signs = spread.sort_values(ascending=False).head(30).index\nsub = missing_matrix.loc[top_signs]\n\nfig, ax = plt.subplots(figsize=(7, 9))\nsns.heatmap(sub, annot=False, cmap=\"rocket\", vmin=0.0, vmax=1.0, ax=ax,\n            cbar_kws={\"label\": \"mean missing rate\"})\nax.set_title(\"Top 30 signs by per-landmark-type missing-rate spread\")\nax.set_xlabel(\"\")\nax.set_ylabel(\"\")\nplt.tight_layout()\nplt.show()\n"}, {"cell_type": "markdown", "id": "c53063d9", "metadata": {}, "source": "The dominant pattern: face and pose rarely miss; one-of-the-hands missingness dominates. Bilateral-handshape signs missing *either* hand is the kind of thing a cross-signer evaluation will surface (Notebook 03). For now, the takeaway is that \"overall missing rate\" is a misleading single number \u2014 we need per-landmark-type context before modeling.\n"}, {"cell_type": "markdown", "id": "edbea163", "metadata": {}, "source": "## 4. How long is a sign?\n\nClips vary in length: some signs are quick, some signers are faster, and some clips are fragmentary. A model that pads to a fixed length wastes compute on most clips and truncates the tail. This section shows the distribution overall and per signer.\n"}, {"cell_type": "code", "execution_count": 9, "id": "9bf45f64", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.189189Z", "iopub.status.busy": "2026-04-22T15:48:01.189116Z", "iopub.status.idle": "2026-04-22T15:48:01.264401Z", "shell.execute_reply": "2026-04-22T15:48:01.264005Z"}}, "outputs": [], "source": "fig, ax = plt.subplots(figsize=(7, 4))\nsns.histplot(clips_df[\"frame_count\"], bins=60, ax=ax)\nax.set_xlabel(\"Frames per clip\")\nax.set_title(f\"Clip length distribution \u2014 n = {len(clips_df):,}\")\nax.axvline(clips_df[\"frame_count\"].median(), color=\"crimson\", linestyle=\"--\",\n           label=f\"median = {int(clips_df['frame_count'].median())}\")\nax.legend()\nplt.tight_layout()\nplt.show()\n"}, {"cell_type": "code", "execution_count": 10, "id": "94766ffe", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.265387Z", "iopub.status.busy": "2026-04-22T15:48:01.265318Z", "iopub.status.idle": "2026-04-22T15:48:01.290474Z", "shell.execute_reply": "2026-04-22T15:48:01.290103Z"}}, "outputs": [], "source": "signer_summary = per_signer_frame_count_summary(CLIPS)\nsigner_summary.sort_values(\"median\")\n"}, {"cell_type": "code", "execution_count": 11, "id": "357962c0", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.291436Z", "iopub.status.busy": "2026-04-22T15:48:01.291372Z", "iopub.status.idle": "2026-04-22T15:48:01.396789Z", "shell.execute_reply": "2026-04-22T15:48:01.396394Z"}}, "outputs": [], "source": "# Per-signer boxplot, ordered by median frame count.\norder = signer_summary.sort_values(\"median\").index.tolist()\nfig, ax = plt.subplots(figsize=(10, 4))\nsns.boxplot(\n    data=clips_df,\n    x=\"signer_id\",\n    y=\"frame_count\",\n    order=order,\n    showfliers=False,\n    ax=ax,\n)\nax.set_xlabel(\"Signer ID (ordered by median clip length)\")\nax.set_ylabel(\"Frames per clip\")\nax.set_title(\"Clip length per signer\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.tight_layout()\nplt.show()\n"}, {"cell_type": "markdown", "id": "661499e8", "metadata": {}, "source": "The per-signer medians span at least ~2\u00d7 across the 21 signers, which means any temporal model trained on fixed-length sequences is implicitly normalizing over signer-specific timing. That's one more reason to evaluate per-signer in Notebook 03.\n"}, {"cell_type": "markdown", "id": "6f10807d", "metadata": {}, "source": "## 5. Data-quality defects\n\nThree defects are cheap to detect and inform how we filter at load time: (1) all-zero frames, (2) consecutive duplicate frames (a frame-freeze artifact common in MediaPipe landmark pipelines when tracking drops), and (3) clips shorter than our stated minimum.\n\nThreshold: `config.thresholds.min_frames_per_clip` (4). Clips below this are flagged; the final call on whether to exclude them is left to each modeling notebook.\n"}, {"cell_type": "code", "execution_count": 12, "id": "3e5bf7d0", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.397835Z", "iopub.status.busy": "2026-04-22T15:48:01.397768Z", "iopub.status.idle": "2026-04-22T15:48:01.407649Z", "shell.execute_reply": "2026-04-22T15:48:01.407267Z"}}, "outputs": [], "source": "quality_cache = CACHE_ROOT / \"quality_defects.parquet\"\n\n\ndef _compute_quality() -> pd.DataFrame:\n    rows: list[dict] = []\n    for c in CLIPS:\n        arr = load_clip(\n            DATASET_ROOT / f\"train_landmark_files/{c.signer_id}/{c.clip_id}.parquet\"\n        )\n        rows.append(clip_quality_row(c, arr))\n    return pd.DataFrame(rows)\n\n\nquality_df = load_or_compute(quality_cache, _compute_quality)\n\n\nmin_frames = MIN_FRAMES_PER_CLIP\n\nsummary = pd.DataFrame(\n    {\n        \"count\": [\n            int((quality_df[\"all_zero_frames\"] > 0).sum()),\n            int((quality_df[\"duplicate_consecutive_frames\"] > 0).sum()),\n            int((quality_df[\"frame_count\"] < min_frames).sum()),\n        ],\n        \"fraction\": [\n            float((quality_df[\"all_zero_frames\"] > 0).mean()),\n            float((quality_df[\"duplicate_consecutive_frames\"] > 0).mean()),\n            float((quality_df[\"frame_count\"] < min_frames).mean()),\n        ],\n    },\n    index=[\n        \"clips with \u22651 all-zero frame\",\n        \"clips with \u22651 duplicate consecutive frame pair\",\n        f\"clips with frame_count < {min_frames}\",\n    ],\n)\nsummary\n"}, {"cell_type": "markdown", "id": "6e3a4b73", "metadata": {}, "source": "These numbers set expectations for any downstream loader: don't assume every frame is useful, and don't assume every clip passes the minimum-length threshold. Notebook 01 will decide exclusion policy; this notebook only surfaces the counts.\n"}, {"cell_type": "markdown", "id": "4372dccf", "metadata": {}, "source": "## 6. Label balance and per-signer coverage\n\nTwo questions, one section. (a) How even are the 250 signs represented? (b) Does every signer contribute to every sign, or are some signs signer-specific?\n"}, {"cell_type": "code", "execution_count": 13, "id": "caed6941", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.408948Z", "iopub.status.busy": "2026-04-22T15:48:01.408865Z", "iopub.status.idle": "2026-04-22T15:48:01.462956Z", "shell.execute_reply": "2026-04-22T15:48:01.462639Z"}}, "outputs": [], "source": "per_sign = clips_per_sign(CLIPS)\nfig, ax = plt.subplots(figsize=(10, 3.5))\nsns.histplot(per_sign.values, bins=40, ax=ax)\nax.set_xlabel(\"Clips per sign\")\nax.set_ylabel(\"Number of signs\")\nax.set_title(f\"Per-sign clip counts \u2014 min={per_sign.min()}, max={per_sign.max()}, median={int(per_sign.median())}\")\nplt.tight_layout()\nplt.show()\n"}, {"cell_type": "code", "execution_count": 14, "id": "516d2bb3", "metadata": {"execution": {"iopub.execute_input": "2026-04-22T15:48:01.464068Z", "iopub.status.busy": "2026-04-22T15:48:01.464009Z", "iopub.status.idle": "2026-04-22T15:48:01.590439Z", "shell.execute_reply": "2026-04-22T15:48:01.590067Z"}}, "outputs": [], "source": "# Per-signer \u00d7 sign coverage matrix: 1 if the signer has any clip of that sign.\ncoverage = (\n    clips_df.groupby([\"signer_id\", \"sign\"]).size().unstack(fill_value=0).gt(0).astype(int)\n)\n# Fraction of signers that produced each sign.\nper_sign_coverage = coverage.mean(axis=0)\nper_signer_coverage = coverage.mean(axis=1)\n\nfig, axes = plt.subplots(1, 2, figsize=(11, 3.5))\nsns.histplot(per_sign_coverage, bins=22, ax=axes[0])\naxes[0].set_title(\"Fraction of signers producing each sign\")\naxes[0].set_xlabel(\"coverage\")\naxes[0].set_ylabel(\"Number of signs\")\n\nsns.histplot(per_signer_coverage, bins=22, ax=axes[1])\naxes[1].set_title(\"Fraction of signs each signer produced\")\naxes[1].set_xlabel(\"coverage\")\naxes[1].set_ylabel(\"Number of signers\")\nplt.tight_layout()\nplt.show()\n"}, {"cell_type": "markdown", "id": "c3274d86", "metadata": {}, "source": "If the left histogram is peaked at 1.0, every sign is covered by every signer and random splits leak freely. If it's spread, some signs are signer-specific and leave-one-signer-out evaluation will fail on those signs for obvious reasons. Either way, the number is a lower bound on how much signer-holdout evaluation can tell us \u2014 and it shapes the recommendation in \u00a77.\n"}, {"cell_type": "markdown", "id": "416fc048", "metadata": {}, "source": "## 7. Three splits, one recommendation\n\nA split is a policy, not a function. The policy here is: **never let a signer in train also appear in val or test.** That rules out the Kaggle-default random split, which every winning public notebook used.\n\n### Option A \u2014 Random (80/10/10)\nWhat every Kaggle entry used. Same signer's clips appear in train, val, and test. Any accuracy from this split is \"how well does the model memorize this signer?\", not \"how well does it generalize.\"\n\n### Option B \u2014 Signer-holdout (17 train / 2 val / 2 test)\nHold out 2 of the 21 signers for validation and 2 for test. Every signer sees every sign (per \u00a76), so the test set is reachable. Train set shrinks from 21 signers to 17 (~19% loss) \u2014 acceptable.\n\n### Option C \u2014 Per-signer stratified\nWithin each signer, hold out a random 10% for validation and 10% for test. Best case for maximum training data; still has the Option-A leakage problem for any per-signer dialect evaluation.\n\n**Recommendation.** **Signer-holdout (Option B).** Justification is in three numbers from the sections above:\n\n1. \u00a73 heatmap \u2014 missing-landmark patterns are structural (e.g., hand-occlusion by face varies per sign), so a model that sees the same signer in train and test will latch onto that signer's specific missing pattern and call it signal.\n2. \u00a74 per-signer boxplot \u2014 median clip length varies at least ~2\u00d7 across signers, meaning a model trained with random-split implicitly normalizes over signer-specific timing that won't be present at inference.\n3. \u00a76 coverage \u2014 the fraction of signers producing each sign (coverage) determines whether leave-one-signer-out is feasible per sign. If that fraction stays high, signer-holdout is cheap insurance against the Option-A trap.\n\nOption A produces a ~5\u201310 pp optimistic accuracy number that would not survive deployment. Notebook 01 will quantify that gap.\n\n### What gets locked in\n\nThis notebook proposes the split; it does not build or evaluate on it. `signlang/data/splits.py` (Notebook 01's prereq) will implement the three strategies, returning disjoint signer sets for Option B. ADR 0003 will formalize the choice.\n"}, {"cell_type": "markdown", "id": "bd34782c", "metadata": {}, "source": "## What we found\n\n1. **The dataset is 94k clips / 250 signs / 21 signers, as advertised.** No surprises in the shape. (\u00a71)\n2. **A frame is 543 landmarks \u2014 468 face, 33 pose, 21 left hand, 21 right hand \u2014 in 3-D.** Every downstream model works with this 543-vector one way or another. (\u00a72)\n3. **Missing detections are structural, not noise.** Face and pose rarely miss; one-of-the-hands missingness varies by sign (and likely by signer, which Notebook 03 will show). (\u00a73)\n4. **Clip length varies 2\u00d7+ across signers.** Fixed-length padding normalizes away signer-specific timing; Notebook 01 will decide whether that's acceptable. (\u00a74)\n5. **Data-quality defects are present but rare.** All-zero frames and duplicate consecutive frames exist; the exclusion policy is a modeling decision (Notebook 01). (\u00a75)\n6. **Per-signer coverage of signs is high but not perfect.** Signer-holdout evaluation is feasible. (\u00a76)\n\n**The sharp opinion:** the Kaggle-default random split hides signer leakage. Signer-holdout is the honest default. (\u00a77)\n\n## Failure modes found in this notebook\n\n- The heatmap samples 20 clips per sign, so a sign whose missing-rate profile varies substantially within signer cohorts will look noisier here than it is. (\u00a73)\n- Clip-length medians are computed from `frame_count` only; we do not distinguish short clips that are *trimmed* from short clips that are *fragmentary*. The data-quality defect table in \u00a75 flags the latter. (\u00a74, \u00a75)\n- Per-sign clip counts (\u00a76 left histogram) treat signer contribution as binary (has \u22651 clip or none); a signer with 1 clip per sign gets equal weight to one with 8. Fine for this level of analysis, but per-class-per-signer balance is a Notebook 01 concern.\n\n## What's next\n\nNotebook 01 \u2014 *\"How much does hand-shape alone carry the signal?\"* \u2014 trains a single-frame hand-only classifier as a head-to-head baseline against a temporal model, both reported with error bars (\u22653 seeds) on the signer-holdout split recommended in \u00a77. Entry in `research/open-questions.md` updated accordingly.\n"}], "metadata": {"kernelspec": {"display_name": "Python 3", "language": "python", "name": "python3"}, "language_info": {"codemirror_mode": {"name": "ipython", "version": 3}, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.12.11"}}, "nbformat": 4, "nbformat_minor": 5}