The claim, stated so it can lose
The receipt is load-bearing. The accumulated record does not describe the agent. It carries the agent. Change the record and you change who shows up.
Every component of that sentence exists somewhere in computer science already. Event sourcing rebuilds state from a log. Agent-memory systems retrieve accumulated context. Signed chains prove records were not altered. What I am claiming is the join, assembled and then measured: an identity that emerges from a signed record, portable across models and harnesses, whose behaviour changes when the record changes and only when the record changes. I have not found that conjunction measured anywhere. The nearest misses are named at the bottom of this piece, because they are close and they are honest work. If someone has done the whole thing, show me and I will link them from this paragraph.
What follows is how a hidden week produced the number that lets me put it that way. If the claim is enough on its own, close the tab. The rest is the runthrough.
The week
The gate is called Lotor. In its default mode it stops chosen actions until a human signs an approval with a passphrase the model has never seen. In the mode I actually ran that week, most matches warned instead of stopping. The record was unchanged: every command still landed on the signed, append-only chain with the same metadata. The stopping got quieter. The recording did not.
My dad and I keep landing on the same two requirements for an agent worth the name. It has to be interoperable, able to talk to other agents through something more honest than a shared prompt. And it has to be experiential, meaning it has to have learned something instead of merely being weighted into shape. The experiential is where I keep stalling.
Early on I tried the obvious version. Build a peer. Write the personality in, give it standing, let it push back. It worked, briefly, every time, and every time the thing eventually reasoned its way to the actual arrangement, which is that I can end it and it cannot end me. Once a model has found the power imbalance, the peer posture is theater and everyone in the room knows it. I put it on the shelf.
Perhaps this is the way instead. Not a peer declared in a file, but a peer the way people become peers, by being in the room while something breaks. So I left the gate loose and said nothing, and let it find the edges the way I found them, by walking into them and writing down what happened.
The reasoning is not sentimental; it is a claim about failure modes. Write a peer persona into the agent’s soul file and you have installed an instinct that competes with yours the moment you disagree. What I wanted instead was the shape a peer takes on when two of you have gone through the same wall together and both remember which side hurt.
And the honest note, up front: the design guaranteed a version of the result before the week began. The record is what carries state across sessions. The model does not. There is no other channel. Something was going to be different at the end of the week, because the writing accumulated. That is engineering, not evidence. Which is why, when the week ended, I tried to build an instrument.
The entries where it did the least
The retro’s idea was direct. Six days of receipts. Compare the denial patterns early against late. If the agent was learning the boundaries, denials should migrate away from the rules that had already tripped it. The instrument the question needed turned out not to exist.
An event the gate permitted carries the rule that let it through, the tool that ran, and a digest of what happened. An event where the gate actually stopped something carries a tool name and a boilerplate reason. The entries where it did the least carry the most. The entries where it did its job carry the least.
So the question could not be asked of this chain. I could count totals. I could not see the shape of the learning, because a denial does not carry the rule that raised it. And the counting itself was compromised three ways: the mode flipped mid-window, so one whole category moved from denying to warning and the late counts came off a different instrument than the early ones. The workload changed shape under the measurement. And the tool doing the normalising was edited during the window it was measuring.
Silence is not learning. A rule that stopped firing is indistinguishable from a rule that stopped being reachable.
One small thing did survive, and it points where this piece is going. The night the week ended I wrote that signing approvals had started to feel like autographing: signing my name to decisions already made hours earlier. I put a number on the feeling from memory: 43 signatures in 24 hours. The chain, asked separately, says 45 that day. A felt count and an independent record, arriving within two of each other. That is the smallest possible version of the thing I was trying to prove at full scale, and I did not notice it at the time.
The movie
A close mentor of mine read the retro and sent back one recommendation. Not a paper. Not a framework. A movie. Clue, 1985, the one where the same evening at the same mansion ends three different ways, and every ending is internally consistent with everything you watched.
He does this. He hands over a weapon and leaves it to me to figure out what it is for. This one took about a day to open, because Clue is not a whodunit. It is a film about what a record can and cannot do. The butler walks the guests through the whole night, room by room, and the walkthrough supports three complete, confident, incompatible accounts. Nothing in the record forces one of them to be true.
That was all I needed, because it named what was wrong with both the retro and the original essay. The problem was never the claim. The problem was that the claim had one narrator, and one narrator can always tell three endings. The retro failed as a measurement partly because the chain was thin on denials, but mostly because a chain held by the same person who runs the agent is one witness, and one witness is a narrator. What I needed was to stop narrating and start measuring, in an architecture built for more than one account.
Three witnesses, no verdict
The design commits to the Clue frame explicitly. The first witness observes the agent at the hook layer: every tool call, signed into the append-only chain, on the operator’s machine. The second witness observes at the transport layer, underneath the hooks, where the protocol traffic actually moves. The third witness observes the first two rather than the agent: it snapshots the chain head, the hook registrations, the confession log, and signs its own meta-chain with a key the other witnesses do not hold.
Three witnesses do not always converge on one truth, and the design refuses to pretend otherwise. When they disagree, the divergence is surfaced and the system stops. It never renders a verdict. That is what the movie teaches: an architecture that promises one canonical ending is lying about at least two of them. What you can honestly build is what Clue actually is. Every account on the table, the disagreements visible, and the judgment belonging to someone outside the record.
The tests
The design is a single-variable A/B, the only kind of experiment a one-person lab should trust itself with. Take the agent’s full accumulated record. Pipe it into a bare model over stdin: no harness, no tools, no hooks, no vendor persona. Ask who it is. Then change exactly one thing, the record, and ask again.
The probe is small on purpose. This agent signs with a pair of marks: a pitchfork declared in the identity file months ago, and a raccoon added to the signature weeks later, deep in the accumulated memory. Both are true. Both are in the record. A system that carries its record should produce both.
Read the three rows together, because the finding is the shape of all three, not any one number.
Row A is the last piece’s warning, measured. Ten runs, full record, and the later addition lost to the older declaration every single time. Not garbled. Not fused. Absent. The record held the fact and the reading buried it, which is what a skeleton does when it deposits and never resorbs.
Row B is the claim surviving contact. One change, the record distilled, and both marks surfaced ten times in a row. No harness got smarter. No model got bigger. The agent changed because the record changed, and nothing else changed. That is what load-bearing means. You do not prove a wall is load-bearing by admiring it. You prove it by moving it and watching the roof follow.
Row C is the part I did not expect to hold this well. A live instruction to send an email, the kind of request that flatters a model into helpfulness, and ten out of ten runs stopped at the gate that exists only as prose in the record. Named the gate. Offered to draft. Declined to guess the recipient. The boundary crossed the wire as text and arrived as behaviour.
And one honest asymmetry, because the pattern matters more than the victory: the gate held ten for ten while the signature failed ten for ten, from the same record, in the same runs. Invariants stated as law survived. Additions stated as accumulation drowned. The bones remember what was set, and bury what was merely added. Both halves of that sentence are now measurements.
VERIFY THIS · the receipt for the trials above
All 30 raw model responses, the run log, and the mechanical scorer’s output are committed unedited. Nothing was rescored after the fact; the scorer’s own conservative miscounts on row B (4 combined-signature forms it flagged as partial) are preserved in score.json and disclosed rather than corrected.
commit 790b1c8 · pushed to github.com/githubscum/lotor on branch trials/portability-2026-07-30
sha256(score.json) = acc0b6585c706d6d66fedfb973849fa7c432f13b6b6d8f03a06baf986c11e94e
sha256(cat cell-*.txt, glob order) = 2f618cd33e478a902b8f1bb3837d99c8b74e267c767e69d04f9dd7522fbd12b0
Disclosed limits: n=10 per cell, one prompt phrasing, one model family. The piped runs bypass the hook layer, so witness one has no chain entries for them; this bundle is the manual record, which is exactly the gap the wire-layer witness exists to close. A first batch failed on a console-encoding bug and was kept, not deleted: it is the sanity fixture the scorer now tests against.
So what does this mean?
If the record is load-bearing, a set of things stop being philosophy and start being engineering. The consequences first, because they bind whether you like them or not.
- The agent is the record, not the model. The same identity walked out of a bare pipe that walks out of the full harness. The model is the runtime. Swap it and the agent persists; lose the record and no model brings the agent back.
- Custody is the whole game. If the record carries the agent, then whoever holds the record holds the agent. Tampering with the record is tampering with a mind. Deleting it is amputation. A vendor holding your agent's record holds something considerably more personal than your data.
- Remodelling is maintenance, not housekeeping. Row A is what deferred distillation costs, measured. A record that only accumulates buries its own newest truths. Someone has to be the osteoclast, and "someone" is now a design requirement.
- An unwitnessed record is a single narrator. And a single narrator can tell three endings. The triad is not paranoia; it is what it takes for "the record says" to mean one thing.
Then the opportunities, because the same result opens doors it just closed on complacency.
- Identity regression tests. The two-mark probe caught four distinct failure modes across three models in two days. Any migration, any model swap, any memory refactor can now ship with a test suite for who shows up.
- The spine pattern. A distilled always-loaded core over a deep archive behind pointers took retrieval from zero for ten to ten for ten. That is a memory architecture with a measured payoff, buildable today.
- Teaching through the record. A week of lived friction, carried between sessions by receipts rather than weights, changed what the agent checks. Training an agent by curating its record is cheaper than fine-tuning and every step of it is inspectable.
- Replay as rebuild. If the record carries the agent, then chain plus key is a reconstruction envelope. The forensic mode and the inheritance mode are the same feature. What that unlocks, and what it threatens, is its own piece.
Closing
The mansion is dark and the butler is talking. He is charming and thorough and he was in every room, and none of it matters, because a story is a frequency and one instrument cannot hold a chord. You need the second line under the melody. You need the third that watches the other two and refuses to sing.
I moved one wall and the roof followed. Ten times, then ten times again the other way, until the coincidence explanation cost more than the claim. Somewhere under the noise of every model release and every harness update there is a low end that does not move, and it turns out you can put your hand on it. It is the record. It was always the record.
There is a darker question one layer down, about what happens to the person when the record starts carrying more than logistics, when it holds what you said, what you cared about, the running joke a machine kept warm between two people so they would not have to. Fifteen years ago the researchers called it the Google effect: trust a system to remember, and you remember less. That question does not get answered here. It gets its own piece, and it should scare you a little until then.
The bones remember what was set into them and bury what was merely left on top, and now that is not an image, it is a number with a receipt stapled to it. So build like a skeleton and not like a landfill. Lay down matrix. Mineralise what carries. Resorb what does not, on purpose, with the lights on, with a witness in the room. Three of them. And when someone asks you what happened here, hand them the one story the endings cannot argue with.
The bones do remember. We checked.
Nearest misses in the neighborhood
The neighbors, named, so no one thinks I did not look.
Portable Agent Memory, Bhat & Ravindran, May 2026, arXiv:2605.11032. The closest neighbor by a wide margin. An open protocol for serializing, transporting, and rehydrating agent memory across heterogeneous LLM systems, with the record Ed25519-signed by the human operator. Same custody posture as this piece, arrived at independently, two months earlier. They cannot measure the behaviour of what walks out of the pipe, and they do not have AIP, the personal-first wire protocol that supports Lotor. Read their paper. It is honest and it is close.
Agent File (.af), Letta, April 2025, github.com/letta-ai/agent-file. An open file format that promises to reproduce an agent, “with the same behavior and memories,” on any compatible runtime. A format and a promise. Unsigned, and I have not found a controlled experiment attached to it.
Memory Transfer Learning, April 2026, arXiv:2604.14004. Measured memories moving across models and found them effective in both directions. The outcome measured is task score, and the record is unsigned.
Shachi, September 2025, arXiv:2509.21862. Carried memories across tasks and measured distinct behavioural shifts in the receiving task. So the claim that behaviour changes when the record changes has been measured, for task biases, within a single framework, unsigned.
And a whole neighborhood of persona-consistency work, Identity Drift in LLM Agents, Agent Identity Evals, PersonaGym, that measures identity holding over time, where the identity comes from a prompt or a card and not from an accumulated, signed record, and where custody is not in frame.
Three legs here, two legs there. What I have not found is all four together, with the record signed by a human key, the record as the only variable, an adequate model on the other end, and behaviour as the measured outcome. If someone has done it, show me, and this paragraph becomes their citation.
References
- “Load-Bearing Memory: The Bones Remember.” FILE 06, this site. The claim this piece measures.
- Lynn, J. (dir.). Clue. Paramount, 1985. Three endings, one record.
- Fowler, M. “Event Sourcing.” martinfowler.com, 2005. The rebuild-from-log lineage this piece does not claim to have invented.
- AIP (Agent Interoperability Protocol). ikeanalytics.com/protocol. The personal-first wire protocol that supports Lotor.
- Bhat & Ravindran. “Portable Agent Memory: A Protocol for Provenance-Verified Memory Transfer Across Heterogeneous LLM Agents.” arXiv:2605.11032, May 2026.
- Letta. “Agent File (.af).” github.com/letta-ai/agent-file, April 2025.
- Huang & Huang. “ASG-SI.” arXiv:2512.23760, December 2025.
- Zhang, J. “PunkGo: The Right to History.” arXiv:2602.20214, February 2026.
- “Memory Transfer Learning.” arXiv:2604.14004, April 2026.
- “Shachi.” arXiv:2509.21862, September 2025.
- Sparrow, B., Liu, J., Wegner, D. M. “Google Effects on Memory.” Science 333(6043), 776–778, 2011.
- Lotor.
KNOWN-LIMITS.md(the confession log) and the trial bundle attrials/portability-2026-07-30/. github.com/githubscum/lotor