{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"15f40edf","cell_type":"markdown","source":"# PIXEL Table II/III Checkpoint Replay\n\nThis notebook evaluates canonical checkpoints on the fixed test splits. Values labelled accepted-paper reference are comparison targets; rows labelled replayed are computed from model predictions in the current session. The Microsoft grayscale checkpoints are protocol-reproduced artifacts whose three per-run test metrics match the accepted-paper records; their SHA-256 hashes preserve that provenance rather than claiming recovered historical bytes.\n","metadata":{}},{"id":"2855bddf","cell_type":"code","source":"from pathlib import Path\nimport subprocess\nimport sys\n\ncandidates = [Path('/kaggle/working/source/reproduce_tables.py')]\ncandidates.extend(Path('/kaggle/input').rglob('reproduce_tables.py'))\nSCRIPT = next((path for path in candidates if path.is_file()), None)\nif SCRIPT is None:\n    raise FileNotFoundError('Attach the PIXEL source Dataset or copy source/ to /kaggle/working/source')\nCONFIG = SCRIPT.parent / 'config' / 'reproduce_kaggle.json'\nprint('Script:', SCRIPT)\nprint('Config:', CONFIG)\n\ndef run(*args, check=True):\n    command = [sys.executable, str(SCRIPT), '--config', str(CONFIG), *args]\n    print('$', ' '.join(command))\n    return subprocess.run(command, check=check)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-21T11:36:01.166293Z","iopub.execute_input":"2026-07-21T11:36:01.167030Z","iopub.status.idle":"2026-07-21T11:36:52.716760Z","shell.execute_reply.started":"2026-07-21T11:36:01.166998Z","shell.execute_reply":"2026-07-21T11:36:52.716179Z"}},"outputs":[],"execution_count":null},{"id":"5d66aa2e","cell_type":"markdown","source":"## 1. Artifact coverage\nThe expected result is 60/60 checkpoints with all four fixed-split JSON files available.\n","metadata":{}},{"id":"641ebafb","cell_type":"code","source":"run('inventory')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-21T11:36:59.908957Z","iopub.execute_input":"2026-07-21T11:36:59.909217Z","iopub.status.idle":"2026-07-21T11:37:00.036395Z","shell.execute_reply.started":"2026-07-21T11:36:59.909194Z","shell.execute_reply":"2026-07-21T11:37:00.035604Z"}},"outputs":[],"execution_count":null},{"id":"0701d2af","cell_type":"markdown","source":"## 2. Accepted-paper reference\nThese values are printed from the machine-readable reference copied from `PIXEL_FINAL.pdf`; this cell does not perform inference.","metadata":{}},{"id":"25b22820","cell_type":"code","source":"run('reference')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-21T11:37:08.423199Z","iopub.execute_input":"2026-07-21T11:37:08.424049Z","iopub.status.idle":"2026-07-21T11:37:08.538666Z","shell.execute_reply.started":"2026-07-21T11:37:08.424000Z","shell.execute_reply":"2026-07-21T11:37:08.537847Z"}},"outputs":[],"execution_count":null},{"id":"f987216b","cell_type":"markdown","source":"## 3. Short checkpoint replay for presentation\nThis evaluates Microsoft R-only and full PPS ResNet-50 run 1, plus EfficientNet-B3 run 1. It demonstrates the same data/model path without spending the video on all 60 checkpoints. Because this mode evaluates one seed, its sample SD is undefined; use the packaged three-seed full-replay summaries for paper-level mean/SD verification.\n","metadata":{}},{"id":"915be9b6","cell_type":"code","source":"RUN_PRESENTATION_REPLAY = True  # Short replay used during rehearsal/recording\nif RUN_PRESENTATION_REPLAY:\n    run(\"evaluate-table3\", \"--datasets\", \"microsoft\", \"--channels\", \"R,RGB\", \"--runs\", \"1\")\n    run(\"evaluate-table2\", \"--datasets\", \"microsoft\", \"--models\", \"eb3\", \"--runs\", \"1\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-21T09:42:38.792140Z","iopub.execute_input":"2026-07-21T09:42:38.792931Z","iopub.status.idle":"2026-07-21T09:45:25.811101Z","shell.execute_reply.started":"2026-07-21T09:42:38.792903Z","shell.execute_reply":"2026-07-21T09:45:25.810187Z"}},"outputs":[],"execution_count":null},{"id":"a48b8cc9","cell_type":"markdown","source":"## 4. Full checkpoint replay\nRun this outside the short video. Table III evaluates all 42 ablation checkpoints. Table II evaluates all model rows, including the three canonical Microsoft grayscale protocol-reproduced checkpoints. Output CSV files are written to `/kaggle/working/pixel_table_replay`. The summary records each row-specific aggregation rule and checks all means and standard deviations against the accepted-paper values at two decimal places; the completed full-replay CSVs are also packaged in `source/results/`.\n","metadata":{}},{"id":"a52a62db","cell_type":"code","source":"RUN_FULL_REPLAY = True  # Keep False while recording; full replay is already audited\nif RUN_FULL_REPLAY:\n    run(\"evaluate-table3\", \"--datasets\", \"malimg,microsoft\", \"--channels\", \"R,G,B,RG,RB,GB,RGB\", \"--runs\", \"1,2,3\")\n    run(\"evaluate-table2\", \"--datasets\", \"malimg,microsoft\", \"--models\", \"resnet_gray,resnet_pps,eb3,eb4\", \"--runs\", \"1,2,3\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-21T09:47:25.281049Z","iopub.execute_input":"2026-07-21T09:47:25.281840Z","iopub.status.idle":"2026-07-21T10:14:56.492453Z","shell.execute_reply.started":"2026-07-21T09:47:25.281809Z","shell.execute_reply":"2026-07-21T10:14:56.490903Z"}},"outputs":[],"execution_count":null}]}