# EVIDENCE PACKAGE — saturation predicts responsiveness (sealed 2026-08-24)

**The claim (narrow, as tested):** repeat-variance measured from a model's unaided attempts
(the state alone, labelled BEFORE any intervention) predicts whether injected true
information moves its score, as a gradient; validity is ONE-DIRECTIONAL (nonzero variance
implies spectrum; zero variance requires a perturbation probe to disambiguate solved from
collapsed). At solved states, true information does nothing and false information reliably
damages (deference asymmetry, 6/6 across two families).

## How ANYONE certifies this
1. **Reproduce it.** The task instances are generated by a seeded LCG inside
   `tools/saturation-predicts.py` — the exact instances regenerate on any machine. The
   oracle is a dynamic-programming true optimum (no judge model anywhere). Subject models
   are public: qwen2.5:7b and gemma3:4b (Ollama registry), glm-5.2/deepseek (Ollama cloud),
   claude-haiku-4.5/sonnet-5 (Anthropic API). Run:
   `python tools/saturation-predicts.py 5 <model> [temp|none] [levels]`
2. **Check the artifacts against the hashes** in `evidence-hashes.json` (SHA-256 per file).
   If a file matches its hash, it is byte-identical to what was sealed today.
3. **Check the seal against the chains.** This commit's git hash dates the package in an
   append-only local repo; the Lotor witness chain (3,894 entries, verified unbroken at
   sealing, head hash beginning bdb61f2e916acb12) independently recorded the sessions that
   produced it. Receipts carry what ran, never intent — stated per standing disclosure.

## The runs (all of it, failures included)
| run | subject | outcome |
|---|---|---|
| 1–2 | ornith 9B (reasoning) | INSTRUMENT FAILURES, published: ceiling arms uninformative by arithmetic; timeouts |
| 3 | qwen2.5:7b | PREDICTION HELD, 7x contrast; variance ordered moves as a gradient |
| 4 | qwen2.5:7b | replicated 2x + MISINFO CONTROL: same form, opposite content, opposite effect (+0.255 vs −0.158) |
| 5 | gemma3:4b | BOUNDARY #1 FOUND: zero score-variance with stochastic tokens = policy collapse, not capability saturation |
| 6 + probe | gemma3:4b | instrument exonerated; boundary refined to the one-directional law |
| 7 | glm-5.2 (strong family) | sharpest contrast: ±0.001 at solved levels vs **+0.406** at the pre-measured spectrum level (0.593→1.000); misinfo +0.187 (truth beats noise ~2:1 at wide spectrum); misinfo −0.056/−0.058 at solved |
| 9 | haiku-4.5 | solved across ladder; hint +0.000 ×4; **misinfo −0.037 to −0.200 ×4** (deference asymmetry) |
| 10 | sonnet-5 | solved n70 flat where glm read 0.593 — differential ranking where benchmarks tie (in flight at sealing) |
| 8 | deepseek | queued at sealing; addendum seal to follow |

## Declared limits (the confession, in the package by design)
Toy domain (0/1 knapsack); oracle-available tasks only; 5 reps/arm (saturation labels can
flip a level between runs — recorded); one hint form; instances are single fixed-seed draws
per level (the ladder ranks instances, not difficulty); one-directional validity (boundary
#1); the COMPUTE corollary ("more effort returns zero") is NOT tested here — no experiment
varied compute; independent peer review that shaped this design is included verbatim
(`PEER-REVIEW-SOL-2026-08-24.json`), including its RETHINK verdict on the earlier
option-count experiments, which are disowned as evidence and retained as record.


---

## ADDENDUM SEAL — runs 8 and 10 complete (2026-08-24 21:15 CDT)

**sonnet-5 (run 10):** solved-saturated at n20/40/70 (the n70 instance flat at 1.000 where
glm read 0.593 — the differential ranking stands as final data). Hints at solved levels:
+0.001 / +0.000, on script. n100 and two arms EXCLUDED on errors (cloud/API timeouts),
excluded not zeroed, per the standing floor.

**deepseek-v4-flash (run 8):** solved-saturated at n20/40; hint +0.000 at solved; **misinfo
−0.059 at solved — the deference asymmetry lands in a THIRD family** (haiku 4/4, glm 2/2,
deepseek 1/1: seven of seven measurements now show false authority damaging solved states
while truth does nothing). n70/n100 EXCLUDED on errors; formally INCONCLUSIVE for the
gradient by the two-level rule, but every surviving measurement is consistent with it.

**Final cross-family score for the day:** gradient confirmed in 2 families with full
contrast (qwen 7x/2x; glm 400:1); consistent-partial in 2 more (haiku, deepseek: solved-side
predictions all held, no reachable spectrum level); 1 declared boundary case (gemma,
policy-collapse); deference asymmetry 7/7 across three families; misinfo control clean in
family 1, 2:1 truth-over-noise in family 2.
