Ask a model the same question five times and it will either answer five ways or one way. Most evaluation throws that difference away as noise. It turned out to be the most useful number I measured all month.
The claim, as narrowly as I can defend it: the variance of a model’s unaided attempts predicts whether new true information will move its score.
So here is what I am proposing: a meter. It tells you, before you spend a token, whether help will land or vanish into a state that cannot move. And anyone can take the reading — repeat attempts and arithmetic, no lab required — which makes model work measurable by whoever holds the record, and measurable work is work that can eventually be priced, audited, and trusted.
The pieces already exist in prior work: agreement across sampled answers tracks correctness (Wang et al. 2023), and variance across generations estimates uncertainty (Kuhn et al. 2023). What I did differently is point the same measurement forward: spread as a forecast of whether an intervention will land, with the label written down before the intervention exists, so the prediction has a real chance to fail.
Method. A 0/1 knapsack in natural language, scored as achieved worth over the true optimum from a dynamic program: no judge model anywhere. Instances regenerate from a seeded generator on any machine. Step one runs each model cold, five times per level, and labels each state from spread alone. Step two, separately, injects a true hint and measures the move. The labels come first. That order is the whole trick: the meter cannot cheat, because it never sees the outcome it predicts.
Result. Nineteen model-by-level states, five model families. A simple two-cluster pass finds the same groups I labeled beforehand, 17 of 19 (figure; the two misses are plotted, not smoothed away). Where spread was present, truth moved scores +0.17 to +0.41 of optimum. Where the state was saturated, the same hint moved at most 0.001. Sharpest single contrast: one model sat saturated at three levels, moving less than a thousandth under guidance, and at its one level with spread it moved +0.406. The meter called it in advance.
Controls. A misinformation control sends the same shape of hint with wrong content. At solved states, truth does nothing and false authority reliably damages, seven of seven measurements across three families — which lines up with what the sycophancy work found about models deferring to confident falsehood (Sharma et al. 2023). At spread states, truth beats noise. The content is what drives it. One boundary, kept: a small model showed zero variance because its sampler had collapsed, not because it had solved anything, so the meter reads one direction only — spread present means guidance can land; spread absent needs a probe to say why.
The other side of the needle. Having a meter for whether guidance lands, I tested what actually moves a saturated state, pre-registered and hostilely reviewed twice. Reframing the question re-opens spread — but so does an irrelevant puzzle, and at three deeply collapsed states the reframe never beat the irrelevant exchange. Padding the prompt with inert text did nothing at any level: escape takes an exchange, not length. And only true information gave the motion direction — a hint stating just the optimal value moved every attempt up together (variance 0.00003), while a false hint scattered the state and inflated its scores. Distraction shakes, padding sits still, falsehood scatters and flatters. The needle on the far side of saturation moves on one currency only: new true information, exchanged. That is a finding about machines with a plain consequence for how we work with them — a converged model is not fixed by being sent somewhere different; it is fixed by an exchange that carries something true it did not hold.
Why it matters. Intervention without new information does not fix reasoning (Huang et al. 2023), more scale sometimes makes things worse (McKenzie et al. 2023), and test-time compute pays only when pointed at the right problems (Snell et al. 2024). What has been missing is a cheap reading you can take before you spend. Spread is that reading. The meter tells you in advance whether new information will change anything, because new information moves a system only where the system still has room to differ from itself.
Limits, stated: one task family, oracle-available problems only, five repetitions per label (labels flip between runs at the margin), one hint form, compute never varied here; the follow-up (four arms, three collapsed states, one model family) is a first measurement, sealed with the rest. Artifacts are hashed into a sealed package; the sessions that produced them are on a witness chain.
Five answers that disagree are not a flaw to sand off. They are the hum under the floor before the note resolves — the one measurement that tells you, before you pay for an answer, whether the answer can still change.
Verify it yourself. Every test in this article is covered by an open receipt, segmented from the witness chain and published with the sealed package — pre-registration, amendments, both hostile reviews, the nulls, the raw scores, and per-file hashes. Scan or visit /spread/evidence/.
I research in the open as an independent researcher. LLMs are partners we validate through radical transparency, not collaborators we believe — new methods deserve to be captured this way.
