{"cells": [{"cell_type": "markdown", "id": "a7926e19", "metadata": {}, "source": "# Student performance: honest nested session CV\n\nThis refresh corrects two misleading parts of the original private draft. The old title said **\u201cNeural Nets\u201d**, but its executable model was TensorFlow Decision Forests; and its shown F1 selected a single global threshold on the same validation labels it scored. This notebook instead compares a question-majority baseline, logistic regression, a clearly labelled aggregate-feature MLP, and compact/full LightGBM models under one nested, session-disjoint protocol.\n\nThe primary claim is deliberately narrow: generalization to unseen gameplay sessions drawn from the same historical collection process. It is not evidence about a new school, cohort, game version, future period, student ability, or causal learning effects.\n\n**Data and rights.** The only source is the official Kaggle competition attachment `predict-student-performance-from-game-play`, credited to Jo Wilder and The Learning Agency Lab. The competition files remain governed by the official [competition rules](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/rules) and [data terms](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data). This notebook neither redistributes them nor asserts a general open-data license. It is a post-competition educational benchmark; the competition closed on 28 June 2023."}, {"cell_type": "markdown", "id": "eba8dbfe", "metadata": {}, "source": "## Evaluation contract\n\n- One prediction row per `(session, question)`; q1\u2013q3 see only levels 0\u20134, q4\u2013q13 only levels 5\u201312, and q14\u2013q18 only levels 13\u201322.\n- Three outer `GroupKFold` splits keep all 18 answers from a session together.\n- Inside each outer-training fold, a second session-disjoint split selects all 18 per-question thresholds.\n- Thresholds are then frozen before the untouched outer fold is scored. No global threshold is tuned on scored labels.\n- The headline metric is pooled binary macro F1, matching the competition definition; mean per-question F1 is not substituted.\n- Calibration diagnostics, question/checkpoint error slices, and 2,000 paired session-bootstrap replicates accompany the point estimates.\n- The MLP is an ordinary 48\u00d724 network over aggregates, **not** a sequence model. LightGBM is labelled as a boosted-tree model."}, {"cell_type": "markdown", "id": "fa4f0233", "metadata": {}, "source": "## Prior work consulted, without copying implementation\n\nThe [first-place write-up](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/writeups/french-touch-1st-place-solution-for-the-predict-st) motivates checkpoint-aware timing/count features and reports that a compact Conv1D was more economical than a Transformer experiment. The [third-place write-up](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/writeups/stablegbt-nn-3rd-place-solution) motivates shared boosted trees with question identity and warns against feature selection outside folds. The [10th-place discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420132) motivates per-question thresholds. These are design references only: no code, prose, predictions, model weights, or output artifacts were copied."}, {"cell_type": "code", "execution_count": null, "id": "eb7ac061", "metadata": {}, "outputs": [], "source": "from pathlib import Path\nimport hashlib, json, os, platform, time, warnings\n\nfrom IPython.display import Image, display\nfrom sklearn.exceptions import ConvergenceWarning\n\nwarnings.filterwarnings(\n    \"ignore\",\n    message=\"Stochastic Optimizer: Maximum iterations.*\",\n    category=ConvergenceWarning,\n)\nwarnings.filterwarnings(\n    \"ignore\",\n    message=\"X does not have valid feature names.*\",\n    category=UserWarning,\n)\n\nHERE = Path.cwd()\nLOCAL_WAVE = HERE.parent if HERE.name == \"candidate\" else HERE\nLOCAL_DATA = LOCAL_WAVE / \"data\"\n\nif (LOCAL_DATA / \"train.csv\").exists() and (LOCAL_DATA / \"train_labels.csv\").exists():\n    # Local evidence run: reuse the retained official competition downloads and\n    # exact benchmark artifacts produced by the same source in this notebook.\n    OFFICIAL_DATA = LOCAL_DATA\n    WORK_ROOT = LOCAL_WAVE\n    EXECUTION_CONTEXT = \"local evidence replay\"\nelse:\n    candidates = [\n        p for p in Path(\"/kaggle/input\").glob(\"predict-student-performance-from-game-play*\")\n        if (p / \"train.csv\").exists() and (p / \"train_labels.csv\").exists()\n    ]\n    if len(candidates) != 1:\n        raise FileNotFoundError(\n            \"Attach the official predict-student-performance-from-game-play competition source.\"\n        )\n    OFFICIAL_DATA = candidates[0]\n    WORK_ROOT = Path(\"/kaggle/working/student-performance-refresh\")\n    WORK_ROOT.mkdir(parents=True, exist_ok=True)\n    EXECUTION_CONTEXT = \"Kaggle clean rebuild\"\n\nprint({\n    \"context\": EXECUTION_CONTEXT,\n    \"source\": \"official predict-student-performance-from-game-play attachment\",\n    \"work_root\": WORK_ROOT.name,\n})"}, {"cell_type": "markdown", "id": "6e0513ae", "metadata": {}, "source": "## 1. Build checkpoint-available features\n\nAggregation is label-free and uses fixed, predeclared event/name/level schemas. Labels are joined only after each session/checkpoint aggregate exists. Session IDs are retained solely for grouping and are asserted absent from model features."}, {"cell_type": "code", "execution_count": null, "id": "642189b6", "metadata": {}, "outputs": [], "source": "\"\"\"Build a label-free, checkpoint-available feature table from official game logs.\n\nThe 4.4 GiB raw event file is scanned one checkpoint at a time so the Kaggle\nCPU runtime never has to retain all high-cardinality strings at once. Every\noutput row is a session/checkpoint aggregate. Labels are joined only after\naggregation and are never used to define schemas or features.\n\"\"\"\n\nfrom __future__ import annotations\n\nimport hashlib\nimport json\nimport os\nimport platform\nimport time\nfrom pathlib import Path\n\nos.environ.setdefault(\"POLARS_MAX_THREADS\", \"2\")\n\nimport polars as pl\n\n\nROOT = WORK_ROOT\nDATA = OFFICIAL_DATA\nOUT = ROOT / \"prepared\"\nOUT.mkdir(parents=True, exist_ok=True)\n\nTRAIN = DATA / \"train.csv\"\nLABELS = DATA / \"train_labels.csv\"\nFEATURES = OUT / \"session_question_features.parquet\"\n\nEVENT_VALUES = [\n    \"checkpoint\",\n    \"cutscene_click\",\n    \"map_click\",\n    \"map_hover\",\n    \"navigate_click\",\n    \"notebook_click\",\n    \"notification_click\",\n    \"object_click\",\n    \"object_hover\",\n    \"observation_click\",\n    \"person_click\",\n]\nNAME_VALUES = [\"basic\", \"close\", \"next\", \"open\", \"prev\", \"undefined\"]\nNUMERIC = [\n    \"page\",\n    \"room_coor_x\",\n    \"room_coor_y\",\n    \"screen_coor_x\",\n    \"screen_coor_y\",\n    \"hover_duration\",\n]\nCARDINALITY_COLUMNS = [\"event_name\", \"name\", \"text\", \"fqid\", \"room_fqid\", \"text_fqid\"]\nLEVEL_GROUPS = [\"0-4\", \"5-12\", \"13-22\"]\nHASH_SEEDS = (20260723, 3342, 9864, 8345)\nQ_TO_GROUP = {\n    **{q: \"0-4\" for q in range(1, 4)},\n    **{q: \"5-12\" for q in range(4, 14)},\n    **{q: \"13-22\" for q in range(14, 19)},\n}\n\n\ndef file_sha256(path: Path, block_size: int = 8 * 1024 * 1024) -> str:\n    digest = hashlib.sha256()\n    with path.open(\"rb\") as stream:\n        while block := stream.read(block_size):\n            digest.update(block)\n    return digest.hexdigest()\n\n\ndef build_features() -> None:\n    started = time.perf_counter()\n    assert TRAIN.exists() and LABELS.exists()\n    pl.Config.set_streaming_chunk_size(50_000)\n\n    columns = [\n        \"session_id\",\n        \"index\",\n        \"elapsed_time\",\n        \"event_name\",\n        \"name\",\n        \"level\",\n        *NUMERIC,\n        \"text\",\n        \"fqid\",\n        \"room_fqid\",\n        \"text_fqid\",\n        \"level_group\",\n    ]\n    events = pl.scan_csv(\n            TRAIN,\n            schema_overrides={\n                \"session_id\": pl.Int64,\n                \"index\": pl.Int32,\n                \"elapsed_time\": pl.Int64,\n                \"level\": pl.Int16,\n                \"page\": pl.Float32,\n                \"room_coor_x\": pl.Float32,\n                \"room_coor_y\": pl.Float32,\n                \"screen_coor_x\": pl.Float32,\n                \"screen_coor_y\": pl.Float32,\n                \"hover_duration\": pl.Float32,\n                \"level_group\": pl.String,\n            },\n            null_values=[\"\"],\n        ).select(columns).with_columns(\n            [\n                pl.col(column)\n                .hash(*HASH_SEEDS)\n                .alias(f\"__{column}_hash\")\n                for column in CARDINALITY_COLUMNS\n            ]\n        )\n\n    expressions: list[pl.Expr] = [\n        pl.len().alias(\"event_count\"),\n        pl.col(\"index\").n_unique().alias(\"unique_index_count\"),\n        pl.col(\"index\").min().alias(\"index_min\"),\n        pl.col(\"index\").max().alias(\"index_max\"),\n        (pl.col(\"index\").max() - pl.col(\"index\").min()).alias(\"index_span\"),\n        pl.col(\"elapsed_time\").min().alias(\"elapsed_min\"),\n        pl.col(\"elapsed_time\").max().alias(\"elapsed_max\"),\n        (pl.col(\"elapsed_time\").max() - pl.col(\"elapsed_time\").min()).alias(\"elapsed_span\"),\n        pl.col(\"__event_name_hash\").n_unique().alias(\"event_name_nunique\"),\n        pl.col(\"__name_hash\").n_unique().alias(\"name_nunique\"),\n        pl.col(\"__text_hash\").n_unique().alias(\"text_nunique\"),\n        pl.col(\"__fqid_hash\").n_unique().alias(\"fqid_nunique\"),\n        pl.col(\"__room_fqid_hash\").n_unique().alias(\"room_fqid_nunique\"),\n        pl.col(\"__text_fqid_hash\").n_unique().alias(\"text_fqid_nunique\"),\n        pl.col(\"level\").n_unique().alias(\"level_nunique\"),\n        pl.col(\"text\").is_not_null().sum().alias(\"text_event_count\"),\n        pl.col(\"fqid\").is_not_null().sum().alias(\"fqid_present_count\"),\n    ]\n    for column in NUMERIC:\n        expressions.extend(\n            [\n                pl.col(column).count().alias(f\"{column}__count\"),\n                pl.col(column).mean().alias(f\"{column}__mean\"),\n                pl.col(column).std().alias(f\"{column}__std\"),\n                pl.col(column).min().alias(f\"{column}__min\"),\n                pl.col(column).max().alias(f\"{column}__max\"),\n            ]\n        )\n    for value in EVENT_VALUES:\n        expressions.append((pl.col(\"event_name\") == value).sum().alias(f\"event__{value}\"))\n    for value in NAME_VALUES:\n        expressions.append((pl.col(\"name\") == value).sum().alias(f\"name__{value}\"))\n    for level in range(23):\n        expressions.append((pl.col(\"level\") == level).sum().alias(f\"level__{level}\"))\n\n    checkpoint_parts = []\n    for level_group in LEVEL_GROUPS:\n        part = (\n            events.filter(pl.col(\"level_group\") == level_group)\n            .group_by(\"session_id\")\n            .agg(expressions)\n            .with_columns(pl.lit(level_group).alias(\"level_group\"))\n            .collect(engine=\"streaming\")\n        )\n        assert part.select(pl.col(\"session_id\").n_unique()).item() == 23_562\n        checkpoint_parts.append(part)\n        print({\"level_group\": level_group, \"checkpoint_rows\": part.height})\n    checkpoint = pl.concat(checkpoint_parts, how=\"vertical\")\n\n    labels = (\n        pl.read_csv(LABELS, schema_overrides={\"session_id\": pl.String, \"correct\": pl.Int8})\n        .with_columns(\n            pl.col(\"session_id\").str.extract(r\"^(\\d+)_q\", 1).cast(pl.Int64).alias(\"base_session_id\"),\n            pl.col(\"session_id\").str.extract(r\"_q(\\d+)$\", 1).cast(pl.Int8).alias(\"q\"),\n        )\n        .drop(\"session_id\")\n    )\n\n    joined = (\n        labels.with_columns(\n            pl.col(\"q\")\n            .replace_strict(Q_TO_GROUP, return_dtype=pl.String)\n            .alias(\"level_group\")\n        )\n        .join(\n            checkpoint,\n            left_on=[\"base_session_id\", \"level_group\"],\n            right_on=[\"session_id\", \"level_group\"],\n            how=\"inner\",\n        )\n        .rename({\"base_session_id\": \"session_id\"})\n        .sort([\"session_id\", \"q\"])\n    )\n\n    assert joined.height == 424_116, joined.height\n    assert joined.select(pl.col(\"session_id\").n_unique()).item() == 23_562\n    assert joined.group_by(\"session_id\").len().select((pl.col(\"len\") == 18).all()).item()\n    assert joined.select(pl.col(\"q\").unique().sort()).to_series().to_list() == list(range(1, 19))\n    assert joined.select(pl.col(\"correct\").unique().sort()).to_series().to_list() == [0, 1]\n    numeric_columns = [name for name, dtype in joined.schema.items() if dtype.is_numeric()]\n    assert joined.select(pl.col(numeric_columns).is_infinite().any()).row(0) == tuple(\n        False for _ in numeric_columns\n    )\n    joined.write_parquet(FEATURES, compression=\"zstd\", statistics=True)\n\n    manifest = {\n        \"created_utc\": time.strftime(\"%Y-%m-%dT%H:%M:%SZ\", time.gmtime()),\n        \"source\": \"Kaggle competition predict-student-performance-from-game-play\",\n        \"source_files\": {\n            \"train.csv\": {\n                \"bytes\": TRAIN.stat().st_size,\n                \"sha256\": file_sha256(TRAIN),\n            },\n            \"train_labels.csv\": {\n                \"bytes\": LABELS.stat().st_size,\n                \"sha256\": file_sha256(LABELS),\n            },\n        },\n        \"output\": {\n            \"path\": FEATURES.name,\n            \"rows\": joined.height,\n            \"columns\": joined.width,\n            \"sessions\": joined.select(pl.col(\"session_id\").n_unique()).item(),\n            \"sha256\": file_sha256(FEATURES),\n        },\n        \"feature_boundary\": \"one session and one available level_group; labels joined after aggregation\",\n        \"label_free_schema\": True,\n        \"memory_contract\": \"three checkpoint-wise streaming scans; 50000-row streaming chunks\",\n        \"cardinality_contract\": {\n            \"method\": \"deterministic 64-bit hash then exact per-group n_unique\",\n            \"columns\": CARDINALITY_COLUMNS,\n            \"hash_seeds\": HASH_SEEDS,\n            \"collision_risk\": \"negligible but non-zero; no raw string is retained by the group aggregate\",\n        },\n        \"sequence_features\": \"index and elapsed-time spans only; no cross-row window state\",\n        \"runtime_seconds\": time.perf_counter() - started,\n        \"environment\": {\n            \"python\": platform.python_version(),\n            \"polars\": pl.__version__,\n            \"platform\": platform.platform(),\n            \"polars_max_threads\": os.environ.get(\"POLARS_MAX_THREADS\"),\n        },\n    }\n    (OUT / \"feature_manifest.json\").write_text(json.dumps(manifest, indent=2) + \"\\n\")\n    print(json.dumps(manifest, indent=2))"}, {"cell_type": "code", "execution_count": null, "id": "9e522c0b", "metadata": {}, "outputs": [], "source": "FEATURES = WORK_ROOT / \"prepared\" / \"session_question_features.parquet\"\nif not FEATURES.exists():\n    build_features()\nelse:\n    print(\"Reusing the exact retained feature table; manifest and source hashes are verified below.\")\n\nfeature_manifest = json.loads((WORK_ROOT / \"prepared\" / \"feature_manifest.json\").read_text())\nassert feature_manifest[\"output\"][\"rows\"] == 424_116\nassert feature_manifest[\"output\"][\"sessions\"] == 23_562\nassert feature_manifest[\"label_free_schema\"] is True\ndisplay(feature_manifest)"}, {"cell_type": "markdown", "id": "82036c28", "metadata": {}, "source": "## 2. Nested benchmark\n\nAll five candidates use the same outer folds. The question-majority predictor is a hard prevalence baseline; logistic regression is the hard linear baseline. Each learned model is fitted once on the inner-fit sessions for threshold calibration and again on the full outer-training sessions for untouched outer-fold evaluation."}, {"cell_type": "code", "execution_count": null, "id": "2634db01", "metadata": {}, "outputs": [], "source": "\"\"\"Nested session-grouped benchmark for the student-performance draft.\n\nThis script compares hard baselines, a compact neural network, and boosted\ntrees on identical outer folds.  Per-question thresholds are chosen only on a\nsession-disjoint calibration split inside each outer-training fold.\n\"\"\"\n\nfrom __future__ import annotations\n\nimport hashlib\nimport json\nimport os\nimport platform\nimport time\nfrom pathlib import Path\n\nos.environ.setdefault(\"OMP_NUM_THREADS\", \"2\")\nos.environ.setdefault(\"OPENBLAS_NUM_THREADS\", \"1\")\nos.environ.setdefault(\"VECLIB_MAXIMUM_THREADS\", \"1\")\n\nimport lightgbm as lgb\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport sklearn\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import brier_score_loss, confusion_matrix, f1_score, log_loss\nfrom sklearn.model_selection import GroupKFold, GroupShuffleSplit\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.pipeline import make_pipeline\nfrom sklearn.preprocessing import StandardScaler\n\n\nROOT = WORK_ROOT\nPREPARED = ROOT / \"prepared\" / \"session_question_features.parquet\"\nOUT = ROOT / \"artifacts\" / \"local-real\"\nOUT.mkdir(parents=True, exist_ok=True)\n\nSEED = 20260723\nN_OUTER = 3\nINNER_CALIBRATION_FRACTION = 0.20\nTHRESHOLD_GRID = np.round(np.arange(0.15, 0.851, 0.02), 2)\nN_BOOTSTRAP = 2_000\n\n\ndef sha256(path: Path) -> str:\n    digest = hashlib.sha256()\n    with path.open(\"rb\") as stream:\n        while block := stream.read(8 * 1024 * 1024):\n            digest.update(block)\n    return digest.hexdigest()\n\n\ndef macro_f1(y: np.ndarray, pred: np.ndarray) -> float:\n    return float(f1_score(y, pred, labels=[0, 1], average=\"macro\", zero_division=0))\n\n\ndef f1_from_counts(counts: np.ndarray) -> float:\n    tn, fp, fn, tp = counts.astype(float)\n    f0 = 2 * tn / max(1.0, 2 * tn + fp + fn)\n    f1 = 2 * tp / max(1.0, 2 * tp + fp + fn)\n    return float(0.5 * (f0 + f1))\n\n\ndef apply_thresholds(q: np.ndarray, probability: np.ndarray, thresholds: dict[int, float]) -> np.ndarray:\n    cuts = np.array([thresholds[int(question)] for question in q], dtype=np.float32)\n    return (probability >= cuts).astype(np.int8)\n\n\ndef optimize_thresholds(q: np.ndarray, y: np.ndarray, probability: np.ndarray) -> tuple[dict[int, float], float]:\n    \"\"\"Deterministic coordinate descent against the pooled official metric.\"\"\"\n    thresholds = {question: 0.5 for question in range(1, 19)}\n    prediction = apply_thresholds(q, probability, thresholds)\n    best = macro_f1(y, prediction)\n    for _ in range(3):\n        changed = False\n        for question in range(1, 19):\n            mask = q == question\n            local_best, local_threshold, local_prediction = best, thresholds[question], prediction[mask]\n            for threshold in THRESHOLD_GRID:\n                trial_local = (probability[mask] >= threshold).astype(np.int8)\n                trial = prediction.copy()\n                trial[mask] = trial_local\n                score = macro_f1(y, trial)\n                if score > local_best + 1e-12 or (\n                    abs(score - local_best) <= 1e-12\n                    and abs(float(threshold) - 0.5) < abs(float(local_threshold) - 0.5)\n                ):\n                    local_best = score\n                    local_threshold = float(threshold)\n                    local_prediction = trial_local\n            if local_threshold != thresholds[question]:\n                thresholds[question] = local_threshold\n                prediction[mask] = local_prediction\n                best = local_best\n                changed = True\n        if not changed:\n            break\n    return thresholds, best\n\n\ndef make_model(name: str, seed: int):\n    if name == \"logistic_full\":\n        return make_pipeline(\n            SimpleImputer(strategy=\"median\"),\n            StandardScaler(),\n            LogisticRegression(C=0.2, max_iter=500, solver=\"lbfgs\", random_state=seed),\n        )\n    if name == \"mlp_full\":\n        return make_pipeline(\n            SimpleImputer(strategy=\"median\"),\n            StandardScaler(),\n            MLPClassifier(\n                hidden_layer_sizes=(48, 24),\n                activation=\"relu\",\n                solver=\"adam\",\n                alpha=0.002,\n                batch_size=2048,\n                learning_rate_init=0.001,\n                max_iter=24,\n                shuffle=True,\n                early_stopping=False,\n                random_state=seed,\n            ),\n        )\n    if name in {\"lightgbm_compact\", \"lightgbm_full\"}:\n        return lgb.LGBMClassifier(\n            objective=\"binary\",\n            n_estimators=360,\n            learning_rate=0.035,\n            num_leaves=31,\n            min_child_samples=100,\n            subsample=0.85,\n            subsample_freq=1,\n            colsample_bytree=0.75,\n            reg_alpha=0.10,\n            reg_lambda=1.0,\n            random_state=seed,\n            n_jobs=2,\n            deterministic=True,\n            force_col_wise=True,\n            verbosity=-1,\n        )\n    raise KeyError(name)\n\n\ndef expected_calibration_error(y: np.ndarray, p: np.ndarray, bins: int = 10) -> float:\n    edges = np.linspace(0, 1, bins + 1)\n    ids = np.clip(np.digitize(p, edges[1:-1], right=False), 0, bins - 1)\n    value = 0.0\n    for bucket in range(bins):\n        mask = ids == bucket\n        if mask.any():\n            value += mask.mean() * abs(float(y[mask].mean()) - float(p[mask].mean()))\n    return float(value)\n\n\ndef session_confusions(frame: pd.DataFrame, prediction: np.ndarray) -> np.ndarray:\n    work = frame[[\"session_id\", \"correct\"]].copy()\n    work[\"prediction\"] = prediction\n    work[\"tn\"] = ((work.correct == 0) & (work.prediction == 0)).astype(np.int16)\n    work[\"fp\"] = ((work.correct == 0) & (work.prediction == 1)).astype(np.int16)\n    work[\"fn\"] = ((work.correct == 1) & (work.prediction == 0)).astype(np.int16)\n    work[\"tp\"] = ((work.correct == 1) & (work.prediction == 1)).astype(np.int16)\n    return work.groupby(\"session_id\", sort=True)[[\"tn\", \"fp\", \"fn\", \"tp\"]].sum().to_numpy(np.int32)\n\n\ndef bootstrap_uncertainty(\n    frame: pd.DataFrame, predictions: dict[str, np.ndarray], reference: str\n) -> tuple[pd.DataFrame, dict[str, list[float]]]:\n    rng = np.random.default_rng(SEED + 999)\n    confusions = {name: session_confusions(frame, pred) for name, pred in predictions.items()}\n    n_sessions = next(iter(confusions.values())).shape[0]\n    values = {name: np.empty(N_BOOTSTRAP, dtype=np.float64) for name in predictions}\n    for replicate in range(N_BOOTSTRAP):\n        sample = rng.integers(0, n_sessions, size=n_sessions)\n        for name, counts in confusions.items():\n            values[name][replicate] = f1_from_counts(counts[sample].sum(axis=0))\n    rows = []\n    reference_values = values[reference]\n    for name, draws in values.items():\n        delta = draws - reference_values\n        rows.append(\n            {\n                \"model\": name,\n                \"bootstrap_mean\": float(draws.mean()),\n                \"score_ci_low\": float(np.quantile(draws, 0.025)),\n                \"score_ci_high\": float(np.quantile(draws, 0.975)),\n                \"delta_vs_reference_mean\": float(delta.mean()),\n                \"delta_ci_low\": float(np.quantile(delta, 0.025)),\n                \"delta_ci_high\": float(np.quantile(delta, 0.975)),\n                \"p_delta_gt_zero\": float((delta > 0).mean()),\n                \"reference\": reference,\n            }\n        )\n    serializable = {name: values[name].round(8).tolist() for name in values}\n    return pd.DataFrame(rows), serializable\n\n\ndef run_nested_benchmark() -> None:\n    started = time.perf_counter()\n    frame = pd.read_parquet(PREPARED)\n    frame[\"level_group\"] = frame[\"level_group\"].astype(str)\n    assert len(frame) == 424_116\n    assert frame.session_id.nunique() == 23_562\n    assert frame.groupby(\"session_id\").size().eq(18).all()\n\n    excluded = {\"session_id\", \"q\", \"level_group\", \"correct\"}\n    full_features = [column for column in frame.columns if column not in excluded]\n    compact_features = [\n        column\n        for column in full_features\n        if not column.startswith((\"event__\", \"name__\", \"level__\"))\n    ]\n    assert \"session_id\" not in full_features\n    assert set(compact_features).issubset(full_features)\n    assert len(full_features) > len(compact_features) >= 30\n\n    # Question identity is categorical, not an ordinal magnitude.\n    question_one_hot = np.eye(18, dtype=np.float32)[frame.q.to_numpy(np.int16) - 1]\n    numeric_full = frame[full_features].to_numpy(np.float32, copy=True)\n    numeric_compact = frame[compact_features].to_numpy(np.float32, copy=True)\n    x_full = np.concatenate([numeric_full, question_one_hot], axis=1)\n    x_compact = np.concatenate([numeric_compact, question_one_hot], axis=1)\n    del numeric_full, numeric_compact, question_one_hot\n    y = frame.correct.to_numpy(np.int8)\n    q = frame.q.to_numpy(np.int8)\n    sessions = frame.session_id.to_numpy(np.int64)\n\n    models = [\"majority_question\", \"logistic_full\", \"mlp_full\", \"lightgbm_compact\", \"lightgbm_full\"]\n    oof_probability = {name: np.full(len(frame), np.nan, dtype=np.float32) for name in models}\n    oof_prediction = {name: np.full(len(frame), -1, dtype=np.int8) for name in models}\n    fold_rows: list[dict] = []\n    threshold_rows: list[dict] = []\n    timing_rows: list[dict] = []\n\n    outer = GroupKFold(n_splits=N_OUTER)\n    for fold, (outer_train, outer_valid) in enumerate(outer.split(x_full, y, groups=sessions), start=1):\n        assert set(sessions[outer_train]).isdisjoint(set(sessions[outer_valid]))\n        unique_train_sessions = np.unique(sessions[outer_train])\n        inner_fit_session_pos, inner_cal_session_pos = next(\n            GroupShuffleSplit(\n                n_splits=1,\n                test_size=INNER_CALIBRATION_FRACTION,\n                random_state=SEED + fold,\n            ).split(unique_train_sessions, groups=unique_train_sessions)\n        )\n        fit_sessions = unique_train_sessions[inner_fit_session_pos]\n        cal_sessions = unique_train_sessions[inner_cal_session_pos]\n        fit_idx = outer_train[np.isin(sessions[outer_train], fit_sessions)]\n        cal_idx = outer_train[np.isin(sessions[outer_train], cal_sessions)]\n        assert set(sessions[fit_idx]).isdisjoint(set(sessions[cal_idx]))\n\n        for model_offset, name in enumerate(models):\n            model_started = time.perf_counter()\n            if name == \"majority_question\":\n                inner_prevalence = {\n                    question: float(y[fit_idx][q[fit_idx] == question].mean()) for question in range(1, 19)\n                }\n                calibration_probability = np.array(\n                    [inner_prevalence[int(question)] for question in q[cal_idx]], dtype=np.float32\n                )\n                thresholds, calibration_f1 = optimize_thresholds(\n                    q[cal_idx], y[cal_idx], calibration_probability\n                )\n                outer_prevalence = {\n                    question: float(y[outer_train][q[outer_train] == question].mean())\n                    for question in range(1, 19)\n                }\n                probability = np.array(\n                    [outer_prevalence[int(question)] for question in q[outer_valid]], dtype=np.float32\n                )\n            else:\n                matrix = x_compact if name == \"lightgbm_compact\" else x_full\n                inner_model = make_model(name, SEED + 100 * fold + model_offset)\n                inner_model.fit(matrix[fit_idx], y[fit_idx])\n                calibration_probability = inner_model.predict_proba(matrix[cal_idx])[:, 1].astype(np.float32)\n                thresholds, calibration_f1 = optimize_thresholds(\n                    q[cal_idx], y[cal_idx], calibration_probability\n                )\n                final_model = make_model(name, SEED + 10_000 + 100 * fold + model_offset)\n                final_model.fit(matrix[outer_train], y[outer_train])\n                probability = final_model.predict_proba(matrix[outer_valid])[:, 1].astype(np.float32)\n\n            prediction = apply_thresholds(q[outer_valid], probability, thresholds)\n            oof_probability[name][outer_valid] = probability\n            oof_prediction[name][outer_valid] = prediction\n            score = macro_f1(y[outer_valid], prediction)\n            elapsed = time.perf_counter() - model_started\n            fold_rows.append(\n                {\n                    \"fold\": fold,\n                    \"model\": name,\n                    \"train_sessions\": len(np.unique(sessions[outer_train])),\n                    \"calibration_sessions\": len(cal_sessions),\n                    \"valid_sessions\": len(np.unique(sessions[outer_valid])),\n                    \"calibration_macro_f1\": calibration_f1,\n                    \"outer_macro_f1\": score,\n                    \"outer_brier\": brier_score_loss(y[outer_valid], probability),\n                    \"outer_log_loss\": log_loss(y[outer_valid], np.clip(probability, 1e-6, 1 - 1e-6)),\n                    \"outer_ece_10bin\": expected_calibration_error(y[outer_valid], probability),\n                    \"runtime_seconds\": elapsed,\n                }\n            )\n            timing_rows.append({\"fold\": fold, \"model\": name, \"seconds\": elapsed})\n            for question, threshold in thresholds.items():\n                threshold_rows.append(\n                    {\"fold\": fold, \"model\": name, \"q\": question, \"threshold\": threshold}\n                )\n            print(f\"fold={fold} model={name} outer_macro_f1={score:.6f} seconds={elapsed:.1f}\", flush=True)\n\n    for name in models:\n        assert np.isfinite(oof_probability[name]).all()\n        assert set(np.unique(oof_prediction[name])).issubset({0, 1})\n\n    fold_metrics = pd.DataFrame(fold_rows)\n    thresholds = pd.DataFrame(threshold_rows)\n    timings = pd.DataFrame(timing_rows)\n    summary_rows = []\n    for name in models:\n        probability = oof_probability[name]\n        prediction = oof_prediction[name]\n        model_folds = fold_metrics[fold_metrics.model == name]\n        summary_rows.append(\n            {\n                \"model\": name,\n                \"nested_oof_macro_f1\": macro_f1(y, prediction),\n                \"fold_mean_macro_f1\": model_folds.outer_macro_f1.mean(),\n                \"fold_std_macro_f1\": model_folds.outer_macro_f1.std(ddof=1),\n                \"brier\": brier_score_loss(y, probability),\n                \"log_loss\": log_loss(y, np.clip(probability, 1e-6, 1 - 1e-6)),\n                \"ece_10bin\": expected_calibration_error(y, probability),\n                \"predicted_positive_rate\": float(prediction.mean()),\n                \"runtime_seconds_total\": timings[timings.model == name].seconds.sum(),\n            }\n        )\n    summary = pd.DataFrame(summary_rows).sort_values(\"nested_oof_macro_f1\", ascending=False)\n\n    # The legacy draft optimized one global cut on the same validation labels\n    # it reported.  This is a fixed, untuned 0.5 comparator only; no threshold\n    # search is ever performed on pooled outer-fold labels.\n    fixed_threshold_rows = []\n    for name in models:\n        fixed_prediction = (oof_probability[name] >= 0.5).astype(np.int8)\n        nested_score = macro_f1(y, oof_prediction[name])\n        fixed_score = macro_f1(y, fixed_prediction)\n        fixed_threshold_rows.append(\n            {\n                \"model\": name,\n                \"nested_inner_calibrated_macro_f1\": nested_score,\n                \"fixed_0_5_macro_f1\": fixed_score,\n                \"nested_minus_fixed_0_5\": nested_score - fixed_score,\n                \"pooled_outer_threshold_search_performed\": False,\n            }\n        )\n    fixed_threshold_ablation = pd.DataFrame(fixed_threshold_rows)\n\n    prediction_frame = frame[[\"session_id\", \"q\", \"level_group\", \"correct\"]].copy()\n    for name in models:\n        prediction_frame[f\"p__{name}\"] = oof_probability[name]\n        prediction_frame[f\"pred__{name}\"] = oof_prediction[name]\n\n    slice_rows = []\n    for name in [\"mlp_full\", \"lightgbm_full\"]:\n        for slice_type, grouper in [(\"question\", \"q\"), (\"checkpoint\", \"level_group\")]:\n            for value, part in prediction_frame.groupby(grouper, sort=True):\n                pred = part[f\"pred__{name}\"].to_numpy(np.int8)\n                truth = part.correct.to_numpy(np.int8)\n                tn, fp, fn, tp = confusion_matrix(truth, pred, labels=[0, 1]).ravel()\n                slice_rows.append(\n                    {\n                        \"model\": name,\n                        \"slice_type\": slice_type,\n                        \"slice\": str(value),\n                        \"n\": len(part),\n                        \"positive_rate\": float(truth.mean()),\n                        \"predicted_positive_rate\": float(pred.mean()),\n                        \"macro_f1\": macro_f1(truth, pred),\n                        \"false_positive_rate\": float(fp / max(1, fp + tn)),\n                        \"false_negative_rate\": float(fn / max(1, fn + tp)),\n                    }\n                )\n    slices = pd.DataFrame(slice_rows)\n\n    uncertainty, bootstrap_draws = bootstrap_uncertainty(\n        frame,\n        {name: oof_prediction[name] for name in models},\n        reference=\"lightgbm_full\",\n    )\n\n    calibration_rows = []\n    for name in models:\n        probabilities = oof_probability[name]\n        bins = np.clip(np.digitize(probabilities, np.linspace(0, 1, 11)[1:-1]), 0, 9)\n        for bucket in range(10):\n            mask = bins == bucket\n            if mask.any():\n                calibration_rows.append(\n                    {\n                        \"model\": name,\n                        \"bin\": bucket,\n                        \"n\": int(mask.sum()),\n                        \"mean_probability\": float(probabilities[mask].mean()),\n                        \"observed_rate\": float(y[mask].mean()),\n                    }\n                )\n    calibration = pd.DataFrame(calibration_rows)\n\n    plt.figure(figsize=(7, 6))\n    for name in [\"logistic_full\", \"mlp_full\", \"lightgbm_full\"]:\n        part = calibration[calibration.model == name]\n        plt.plot(part.mean_probability, part.observed_rate, marker=\"o\", label=name)\n    plt.plot([0, 1], [0, 1], \"--\", color=\"black\", linewidth=1, label=\"perfect\")\n    plt.xlabel(\"Mean predicted probability\")\n    plt.ylabel(\"Observed correctness rate\")\n    plt.title(\"Nested out-of-fold reliability (descriptive)\")\n    plt.legend()\n    plt.tight_layout()\n    plt.savefig(OUT / \"calibration_reliability.png\", dpi=150)\n    plt.close()\n\n    fold_metrics.to_csv(OUT / \"fold_metrics.csv\", index=False)\n    summary.to_csv(OUT / \"benchmark_summary.csv\", index=False)\n    fixed_threshold_ablation.to_csv(OUT / \"fixed_threshold_ablation.csv\", index=False)\n    thresholds.to_csv(OUT / \"inner_thresholds.csv\", index=False)\n    timings.to_csv(OUT / \"model_timings.csv\", index=False)\n    slices.to_csv(OUT / \"error_slices.csv\", index=False)\n    uncertainty.to_csv(OUT / \"bootstrap_uncertainty.csv\", index=False)\n    calibration.to_csv(OUT / \"calibration_bins.csv\", index=False)\n    prediction_frame.to_parquet(OUT / \"nested_oof_predictions.parquet\", index=False)\n    (OUT / \"bootstrap_draws.json\").write_text(json.dumps(bootstrap_draws) + \"\\n\")\n\n    threshold_stability = (\n        thresholds.groupby([\"model\", \"q\"]).threshold.agg([\"min\", \"max\", \"mean\", \"std\"]).reset_index()\n    )\n    threshold_stability.to_csv(OUT / \"threshold_stability.csv\", index=False)\n\n    protocol = {\n        \"unit_of_generalization\": \"unseen gameplay session from the same collection process\",\n        \"rows_per_session\": 18,\n        \"outer_validation\": f\"{N_OUTER}-fold GroupKFold by session\",\n        \"inner_calibration\": \"GroupShuffleSplit by session inside each outer-training fold\",\n        \"thresholds\": \"18 per-question thresholds optimized only on inner calibration labels\",\n        \"metric\": \"pooled binary macro F1 across all session-question rows\",\n        \"probability_diagnostics\": [\"Brier score\", \"log loss\", \"10-bin ECE\", \"reliability table\"],\n        \"uncertainty\": f\"{N_BOOTSTRAP} paired session bootstrap replicates\",\n        \"neural_model\": \"fixed 48x24 MLP on aggregate features; not a sequence model\",\n        \"no_supervised_feature_selection\": True,\n        \"future_checkpoint_protection\": \"q1-3 use 0-4, q4-13 use 5-12, q14-18 use 13-22\",\n        \"scope_limit\": \"does not test another school, cohort, game version, or future time period\",\n    }\n    (OUT / \"protocol.json\").write_text(json.dumps(protocol, indent=2) + \"\\n\")\n\n    runtime = {\n        \"created_utc\": time.strftime(\"%Y-%m-%dT%H:%M:%SZ\", time.gmtime()),\n        \"wall_seconds\": time.perf_counter() - started,\n        \"python\": platform.python_version(),\n        \"platform\": platform.platform(),\n        \"numpy\": np.__version__,\n        \"pandas\": pd.__version__,\n        \"scikit_learn\": sklearn.__version__,\n        \"lightgbm\": lgb.__version__,\n        \"omp_num_threads\": os.environ.get(\"OMP_NUM_THREADS\"),\n    }\n    (OUT / \"runtime.json\").write_text(json.dumps(runtime, indent=2) + \"\\n\")\n\n    validation_receipt = {\n        \"rows\": int(len(frame)),\n        \"sessions\": int(frame.session_id.nunique()),\n        \"rows_per_session_exactly_18\": bool(frame.groupby(\"session_id\").size().eq(18).all()),\n        \"outer_folds\": N_OUTER,\n        \"models\": models,\n        \"fold_metric_rows\": int(len(fold_metrics)),\n        \"threshold_rows\": int(len(thresholds)),\n        \"question_thresholds_per_model_fold\": bool(\n            thresholds.groupby([\"model\", \"fold\"]).size().eq(18).all()\n        ),\n        \"paired_session_bootstrap_replicates\": N_BOOTSTRAP,\n        \"bootstrap_models\": int(len(uncertainty)),\n        \"error_slice_rows\": int(len(slices)),\n        \"all_oof_probabilities_finite\": bool(\n            all(np.isfinite(oof_probability[name]).all() for name in models)\n        ),\n        \"all_oof_predictions_binary\": bool(\n            all(set(np.unique(oof_prediction[name])).issubset({0, 1}) for name in models)\n        ),\n        \"pooled_outer_threshold_search_performed\": False,\n        \"session_identifier_used_as_feature\": False,\n        \"raw_competition_files_in_artifacts\": False,\n    }\n    assert validation_receipt[\"fold_metric_rows\"] == N_OUTER * len(models)\n    assert validation_receipt[\"threshold_rows\"] == N_OUTER * len(models) * 18\n    assert validation_receipt[\"question_thresholds_per_model_fold\"]\n    assert validation_receipt[\"rows_per_session_exactly_18\"]\n    (OUT / \"validation_receipt.json\").write_text(\n        json.dumps(validation_receipt, indent=2) + \"\\n\"\n    )\n\n    artifacts = []\n    for path in sorted(OUT.iterdir()):\n        if path.name == \"artifact_manifest.csv\" or not path.is_file():\n            continue\n        artifacts.append({\"file\": path.name, \"bytes\": path.stat().st_size, \"sha256\": sha256(path)})\n    pd.DataFrame(artifacts).to_csv(OUT / \"artifact_manifest.csv\", index=False)\n    print(\"\\nFINAL SUMMARY\")\n    print(summary.to_string(index=False))\n    print(\"\\nUNCERTAINTY\")\n    print(uncertainty.to_string(index=False))\n    print(json.dumps(runtime, indent=2))"}, {"cell_type": "code", "execution_count": null, "id": "7fa78fa5", "metadata": {}, "outputs": [], "source": "ARTIFACTS = WORK_ROOT / \"artifacts\" / \"local-real\"\nrequired = {\n    \"artifact_manifest.csv\", \"benchmark_summary.csv\", \"bootstrap_uncertainty.csv\",\n    \"calibration_bins.csv\", \"calibration_reliability.png\", \"error_slices.csv\",\n    \"fixed_threshold_ablation.csv\",\n    \"fold_metrics.csv\", \"inner_thresholds.csv\", \"nested_oof_predictions.parquet\",\n    \"protocol.json\", \"runtime.json\", \"threshold_stability.csv\", \"validation_receipt.json\",\n}\nif not required.issubset({p.name for p in ARTIFACTS.glob(\"*\")}):\n    run_nested_benchmark()\nelse:\n    print(\"Reusing exact audited benchmark artifacts generated by the source above.\")\n\nmanifest = pd.read_csv(ARTIFACTS / \"artifact_manifest.csv\")\nfor row in manifest.itertuples(index=False):\n    path = ARTIFACTS / row.file\n    assert path.exists()\n    digest = hashlib.sha256(path.read_bytes()).hexdigest()\n    assert digest == row.sha256, (row.file, digest, row.sha256)\nprint(f\"Verified {len(manifest)} artifact hashes.\")"}, {"cell_type": "markdown", "id": "3ca872e5", "metadata": {}, "source": "## 3. Results and paired uncertainty"}, {"cell_type": "code", "execution_count": null, "id": "d8951d4d", "metadata": {}, "outputs": [], "source": "summary = pd.read_csv(ARTIFACTS / \"benchmark_summary.csv\")\nfolds = pd.read_csv(ARTIFACTS / \"fold_metrics.csv\")\nuncertainty = pd.read_csv(ARTIFACTS / \"bootstrap_uncertainty.csv\")\nthresholds = pd.read_csv(ARTIFACTS / \"threshold_stability.csv\")\nfixed_threshold = pd.read_csv(ARTIFACTS / \"fixed_threshold_ablation.csv\")\n\ndisplay(summary.round(5))\ndisplay(folds.pivot(index=\"fold\", columns=\"model\", values=\"outer_macro_f1\").round(5))\ndisplay(uncertainty.round(5))\ndisplay(fixed_threshold.round(5))\n\nbest = summary.iloc[0]\nreference = summary.query(\"model == 'lightgbm_full'\").iloc[0]\nprint({\n    \"best_model\": best.model,\n    \"best_nested_oof_macro_f1\": round(float(best.nested_oof_macro_f1), 6),\n    \"lightgbm_full_nested_oof_macro_f1\": round(float(reference.nested_oof_macro_f1), 6),\n})"}, {"cell_type": "markdown", "id": "eb563ad2", "metadata": {}, "source": "The bootstrap resamples complete sessions, so all 18 paired outcomes travel together. A model should not be called better merely because its point estimate is larger; the paired delta interval must support that claim. Fold dispersion remains visible because three folds cannot diagnose every dataset shift."}, {"cell_type": "markdown", "id": "8b24a912", "metadata": {}, "source": "## 4. Calibration and failure slices"}, {"cell_type": "code", "execution_count": null, "id": "f3fa5d96", "metadata": {}, "outputs": [], "source": "display(Image(filename=str(ARTIFACTS / \"calibration_reliability.png\")))\n\nslices = pd.read_csv(ARTIFACTS / \"error_slices.csv\")\nworst_questions = (\n    slices.query(\"slice_type == 'question'\")\n    .sort_values([\"model\", \"macro_f1\"])\n    .groupby(\"model\", as_index=False)\n    .head(5)\n)\ndisplay(worst_questions.round(4))\ndisplay(slices.query(\"slice_type == 'checkpoint'\").round(4))"}, {"cell_type": "markdown", "id": "c48ab3e7", "metadata": {}, "source": "## 5. Threshold stability and leakage audit"}, {"cell_type": "code", "execution_count": null, "id": "a8dbe530", "metadata": {}, "outputs": [], "source": "display(thresholds.sort_values([\"model\", \"q\"]).round(3))\n\nprotocol = json.loads((ARTIFACTS / \"protocol.json\").read_text())\nreceipt = json.loads((ARTIFACTS / \"validation_receipt.json\").read_text())\nassert protocol[\"outer_validation\"].startswith(\"3-fold GroupKFold\")\nassert protocol[\"no_supervised_feature_selection\"] is True\nassert protocol[\"thresholds\"].startswith(\"18 per-question\")\nassert protocol[\"uncertainty\"].startswith(\"2000 paired session\")\ndisplay(protocol)\ndisplay(receipt)"}, {"cell_type": "code", "execution_count": null, "id": "7246fedd", "metadata": {}, "outputs": [], "source": "# Kaggle can pre-create an empty submission placeholder for competition notebooks.\n# Kaggle owns this path and may prohibit deleting it. Record the platform artifact\n# literally; this notebook never writes a submission or calls the submission API.\nplaceholder = Path(\"/kaggle/working/submission.csv\")\nplaceholder_bytes = placeholder.stat().st_size if placeholder.exists() else None\nprint({\n    \"competition_submission_made\": False,\n    \"submission_written_by_notebook\": False,\n    \"platform_owned_submission_placeholder_bytes\": placeholder_bytes,\n})"}, {"cell_type": "markdown", "id": "99bdc945", "metadata": {}, "source": "## Interpretation and limits\n\nThis benchmark is publishable only as a transparent historical validation study. Its usefulness is the protocol, the hard baselines, and the negative/positive comparisons\u2014not a claim of global state of the art. Generic aggregates discard detailed story progression and event order; the MLP is not an event-sequence encoder; session-grouped CV cannot establish transfer to another population; and gameplay correlations are not causal explanations of learning. The fixed-budget MLP reaches its 24-iteration cap and remains a conservative computational baseline rather than a fully tuned neural winner.\n\nAny future model\u2014richer room/object timing features, leakage-safe prior-question meta-features, or a compact Conv1D\u2014must use the same nested session split and paired uncertainty. The competition is closed, so this notebook must not promise or manufacture a new leaderboard submission."}], "metadata": {"kernelspec": {"display_name": "Python 3", "language": "python", "name": "python3"}, "language_info": {"name": "python", "version": "3.11"}}, "nbformat": 4, "nbformat_minor": 5}