# Saturation predicts responsiveness: PREDICTION HELD (2026-08-24, run 3, non-circular)

Model: qwen2.5:7b (pulled for this run -- Isaac's "negative scaling into signal": the weak
model IS the spectrometer; ornith, a reasoning model, either aced or timed out and could
never expose the band). Oracle: computed DP optimum. 5 reps/arm. Labels fixed BEFORE step 2.

| level | base mean | variance | label | hint MOVE |
|---|---|---|---|---|
| anchor n=12 | 0.501 | 0.01116 | spectrum | +0.169 |
| mid16 | 0.322 | 0.01161 | spectrum | +0.191 |
| mid20 | 0.264 | 0.00126 | spectrum | +0.095 |
| mid30 | 0.261 | 0.00041 | SATURATED | -0.021 |

**|move| spectrum = 0.152 vs saturated = 0.021 (7x). PREDICTION HELD** -- new information
moves the score where spectrum was pre-measured and barely moves it where the state was
already saturated. Non-circular: saturation labelled from the state alone, before the
exchange; oracle mechanical; instruction shape identical across arms.

**Stronger than predicted: variance ORDERS the moves** (gradient, not threshold):
0.011->+0.17/+0.19, 0.0013->+0.10, 0.0004->-0.02.

**Bonus 1 -- floor vs ceiling saturation distinguished, answering scout 2's open question in
the toy domain:** zero variance at mean 0.261 = collapsed-on-wrong-basin, NOT solved.
Variance alone cannot tell them apart; variance + oracle-anchored mean can. Done live.
**Bonus 2:** the hint slightly HURT at the floor (-0.021): no spectrum to act on.

**Honest bounds:** one model, one task family, 5 reps, one hint form, single floor level.
A supported prediction in a toy domain, not a law. Next: second model family, second task
family, more floor levels, vary hint size.

**Run history:** run 1 (ornith) inconclusive -- ceiling arms uninformative by arithmetic +
timeouts; run 2 (ornith, 400s) same; run 3 (qwen2.5:7b) landed in ~5 min. Two instrument
failures, one honest kill, one clean result. The instrument lesson: a saturated instrument
cannot measure saturation -- descending capability is descending INTO the signal band.


---

# Runs 4-5 (same day): the misinfo control HELD, and the cross-family arm found the
# instrument's first declared failure boundary

## Run 4 — qwen2.5:7b, replication + MISINFORMATION control
Main effect replicated qualitatively (|move| 0.154 spectrum vs 0.075 saturated — 2x, weaker
than v3's 7x; and the "saturated" label FLIPPED level between runs, so 5-rep variance labels
are themselves noisy — both honest caveats for the write-up).

**The misinfo control killed the demand-characteristics attack:** identical hint form, wrong
content -> at the most-spectrum level, real info +0.255 vs misinfo -0.158. **Opposite content,
opposite effect, same form: CONTENT drives it.** At the saturated level, misinfo did ~nothing
(+0.012), as predicted.

## Run 5 — gemma3:4b: DECLARED FAILURE BOUNDARY #1
Gemma returned **variance 0.00000 at every level** (deterministic decoding) -> everything
labelled saturated -> inconclusive by rule. But hints moved it anyway, and **misinformation
helped MORE than real information** (+0.310 vs +0.038 at mid16): a model collapsed onto a bad
basin is knocked UPWARD by any perturbation, including wrong ones.

**The boundary, declared:** repeat-variance conflates two different zeros —
**capability-saturation** (qwen's floor: unresponsive; misinfo hurts) and **decoder-collapse**
(gemma: the sampler is peaked, not the knowledge; perturbable by anything). **The instrument
is valid only for stochastic samplers.** Scout 2's literature caveat ("policy entropy flat
while perceptual diversity collapses"; agreement manufactured by post-training) hit us within
hours of being filed. Fix in flight: forced-temperature re-measurement (run 6).

**Anomaly logged, not explained:** misinfo at gemma mid16 (+0.310) beat real info (+0.038).
Perturbation-escape from a bad basin is the candidate story; needs its own test before it is
a claim.


---

# Run 6 + the probe: the boundary refined to a ONE-DIRECTIONAL law (2026-08-24 ~17:00)

Temperature 1.0 did NOT restore score-variance (baselines byte-identical to run 5), yet the
hinted arms differed between runs -- internally contradictory, so the instrument became the
first suspect. **Direct probe exonerated it:** gemma is fully stochastic in language
(different animals on every identical call, with and without temperature), while its knapsack
scores were identical 5/5, twice.

**The real phenomenon: gemma3:4b is stochastic in TOKENS and deterministic in DECISIONS.**
It picks the same item-set every time regardless of surface variation. The collapse is at the
policy level, not the decoder -- the inverse of scout 2's cited caveat (policy entropy flat
while perceptual diversity collapses): here token diversity is present while decision
diversity is collapsed.

## Boundary #1, final form (one-directional validity)
- **Nonzero score-variance => spectrum exists.** Reliable in every run.
- **Zero score-variance =/=> no spectrum.** Policy collapse can mask latent capability:
  gemma's +0.377 hint response proves the spectrum existed beneath the collapsed policy.
- **Disambiguator: the perturbation probe.** At zero variance, inject and watch. Response =
  masked spectrum; no response = true saturation (qwen's mid30 floor: -0.021 real, -0.081
  misinfo). The hint arm is not just a test of the theory; it is the second instrument the
  first one needs.

Run ledger for the day: 2 instrument failures (ornith, published as such) -> 1 clean positive
(qwen, 7x) -> 1 replication with the misinfo control (content-driven confirmed, 2x, label
noise disclosed) -> 1 boundary discovery (gemma decision-collapse) -> 1 probe that exonerated
the instrument and refined the boundary to a directional law. Six runs, everything dated,
failures included.
