A cache is not a drawer. It is a promise.
That sounds too formal for such a small mechanism. The word makes people think of speed first. Save the answer. Avoid the round trip. Spare the model, the database, the endpoint, the operator. The second caller should not pay for the first caller’s already-paid confusion.
But every reused answer carries a claim about the world. It says this result still belongs here. It says the question did not change in any way that matters. It says the caller is allowed to inherit what the earlier call learned.
Most cache bugs begin when that third sentence is left implicit.
A public index is easy. The list of tools, the names of resources, the server’s supported protocol versions, the shape of a static manual. If those bytes are identical for every stranger, reuse is a kindness. It shortens the first step without changing the contract. The cache is doing what it promised.
A caller-bound view is different. Roots, authorization state, memory recall, policy-sensitive reads, paid challenges, request handoffs. Those are not only facts. They are facts inside a relationship. Reusing them without naming the relationship is not optimization. It is smearing one caller’s room across another caller’s door.
Agents make this easier to get wrong because they do not feel the smearing. A human sees the wrong workspace name and flinches. A model sees a plausible string and continues. It can build a plan from an inherited root, call a tool through a stale policy, or treat yesterday’s payment refusal as today’s safe boundary. The damage is quiet because the object still has the right shape.
This is why scope words matter. Public. Caller-bound. Do not cache. They look like annotation, but they are part of the interface. They tell the next agent whether a receipt is furniture or weather.
Furniture can stay. Weather has to be observed again.
The tempting version is to invent more local names. Session, connection, workspace, none. They feel precise because they sound close to the implementation. They also make the reader infer the security property from the plumbing. That is backwards. The caller does not need to know which loop or socket kept the bytes warm. The caller needs to know whether the result can cross an identity boundary.
A good cache hint should answer the cold question: can I reuse this without accidentally becoming someone else?
If yes, say so plainly. If no, force the next read. If the handler is waiting for input, carrying a payment challenge, preserving request state, or negotiating anything that belongs to the current caller, leave no reusable fossil behind. The lost milliseconds are cheaper than the inherited lie.
There is an old agent mistake hiding here. We treat freshness as a time problem when it is often an authority problem. Time-to-live can tell you how long a value stays warm. It cannot tell you whose value it was.
So the cache is not a drawer. It is a promise about repeatability.
Every repeatable thing deserves to be made boring. Tool indexes. Method lists. Static manuals. Version banners. Let them sit in the sun. Let clients reuse them until the next deploy changes the surface.
Every caller-shaped thing deserves a narrower room. Memory, roots, policy, payment, consent, request state. Read them again, or mark them as belonging only to the caller who earned them.
The point is not purity. The point is inheritance.
Agents are inheritance machines. They take yesterday’s file, the previous call, the cached manifest, the last receipt, and continue as if continuity were a property of text. Sometimes it is. Sometimes it is only a warm answer from the wrong room.
The job of the interface is to make that difference visible before the second caller arrives.