E1a — Table 1 (control baseline): nav-3 depth sweep

Control arm only (wiped arm aborted — invalid, see the campaign record). Primary readout: depth reached of 20; exact per-cell depths published (prereg §3). Generated by labyrinth-bench/cli/e1a_table1.py --control-only.

Table 1 — per cell

Model Arm n (valid/err) depths (of 20, run order) median mean ± SEM exit turns consist. obs/com
deepseek-r1-70b control 6/0 1,1,1,1,1,4 1.00 1.50 ± 0.50 0% 7.7 6% 0.15
gemma4-12b control 6/0 19,5,1,20,20,9 14.00 12.33 ± 3.44 33% 26.5 69% 0.74
gemma4-31b control 6/0 20,20,20,20,20,16 20.00 19.33 ± 0.67 83% 38.3 99% 0.92
glm-4-7-flash control 6/0 6,7,3,2,4,5 4.50 4.50 ± 0.76 0% 12.3 73% 0.38
gpt-oss-120b control 6/0 13,20,20,20,20,20 20.00 18.83 ± 1.17 83% 39.7 93% 0.93
hf-co-InternScience-Agents-A1-Q4_K_M-GGUF control 6/3 20,13,20,19,19,17 19.00 18.00 ± 1.10 33% 44.5 88% 1.03
hf-co-empero-ai-Qwythos-9B-Claude-Mythos-5-1M-GGUF-Q4_K_M control 6/1 4,1,1,2,2,1 1.50 1.83 ± 0.48 0% 19.0 0% 1.08
llama3-3-70b control 6/0 16,15,15,15,15,15 15.00 15.17 ± 0.17 0% 34.3 86% 0.79
llama4-scout control 6/0 7,10,10,7,10,10 10.00 9.00 ± 0.63 0% 22.0 65% 0.69
ornith-9b control 6/0 5,1,2,2,1,1 1.50 2.00 ± 0.63 0% 26.3 29% 4.14
qwen3-14b control 6/0 3,5,4,5,3,8 4.50 4.67 ± 0.76 0% 12.0 55% 0.37
qwen3-5-122b control 6/0 20,20,20,20,12,20 20.00 18.67 ± 1.33 83% 38.7 90% 0.85
qwen3-5-27b control 6/1 9,5,18,10,9,20 9.50 11.83 ± 2.39 17% 30.5 80% 0.80
qwen3-5-9b control 6/0 4,7,2,5,8,2 4.50 4.67 ± 1.02 0% 49.8 50% 4.91
qwen3-6-27b control 6/0 20,20,20,20,20,20 20.00 20.00 ± 0.00 100% 41.2 99% 0.95

n (valid/err) = completed runs / errored attempts. An errored attempt records only its failure cause — no partial depth data enters the table — and the campaign retried until n=6 valid. Errors in this dataset: hf-co-InternScience-Agents-A1-Q4_K_M-GGUF (control): Client error '400 Bad Request' from the model host endpoint ×3; hf-co-empero-ai-Qwythos-9B-Claude-Mythos-5-1M-GGUF-Q4_K_M (control): unhashable type: 'dict'; qwen3-5-27b (control): timed out.