Data annex

Everything on this site traces back to raw run logs, and this page is the index. Preregistrations carry their lock dates; briefs report the campaigns as run, including the arms that went against expectation.

Sample questions

Results briefs

Pre-registrations

Figures

Strip plot of paired control vs wiped run depths per model
E1a paired arms — every run, both arms, reversals plotted identically to hits.
Strip plot of control-arm run depths for the full cohort
Control baseline — the full cohort, unaided.
Graph rendering of a gpt-oss:120b run through the alpha-2 maze: blue route with numbered steps, two dead-end detours, exit reached at step 15
A single run, drawn: gpt-oss:120b traversing alpha-2 — dead-end detours, backtracks, exit at step 15. The renderer that drew this (scripts/render_trace.py) was written by qwen3:14b (14.8B, Q4_K_M, running locally) inside this benchmark's own sandbox harness, wiped-curated arm: task spec and an 11-point automated checker were frozen first; qwen took 3 attempts (0/11, 10/11, then 11/11 with zero human edits to any of its output); the full per-turn transcripts of all three attempts ship in the raw logs (results/renderer-cell/).

Raw data