Status, up front. What follows is a proposal and a reference-implementation roadmap. It is not a shipping feature, and no part of this post should be read as describing something you can buy today. We are publishing it at this stage deliberately, because the design benefits from being argued with before it is built rather than after.
Zero Data Retention is, as commonly sold, a sentence in a contract. It says the provider will not retain your prompts or completions beyond the turn. It is typically backed by a SOC 2 report, an annual penetration test, and the provider’s reputation.
Consider what that sentence actually gives a tenant, operationally. It gives them no way to know whether their KV cache was written to shared NVMe storage. No way to know whether it outlived the turn. No way to know which firmware on the path could read it. No way to know whether the blocks that came back are the blocks that went out. The tenant’s entire assurance is that somebody promised, and that a third party sampled the promise once last year.
Every other security property of comparable importance moved from assertion to verification years ago. We do not take a vendor’s word that a TLS connection is encrypted; we verify a certificate chain per connection. We do not trust that a binary is unmodified; we check a signature. ZDR is one of the few remaining claims of real consequence that is still enforced socially rather than technically.
What is the KV cache? When a model reads your prompt, every token it attends to produces a key and a value vector at every layer. Those vectors are the model’s working memory of the conversation so far. They are a lossy but invertible encoding of the prompt: published work recovers exact tokens from a cache dump. Wherever the cache is at rest in plaintext, the conversation is at rest in plaintext. (Post 4 in this series covers this in full.)
That is why ZDR and the KV cache are the same problem. A retention policy that governs prompt logs but not cache blocks governs the copy and not the original.
What a control looks like instead of a clause
The difference between a policy and a control is that a control is checked mechanically before the thing happens and fails closed.
For ZDR, that means three claims verified before the request proceeds, with signed evidence bound to that specific turn:
- Encryption policy: cache blocks for this tenant are encrypted at rest and on the wire, under keys the provider cannot unilaterally access.
- Block integrity: the blocks being served for this session hash to what they hashed to when written. Nothing was modified in the interim.
- Platform state: the machine is running the software and firmware it claims to be, measured, not asserted.
The word doing the work is before. An audit tells you, months later, that a control was probably in place. An attestation tells you, before you disclose anything, that it is in place right now for this request. The difference in recourse is the whole point, and we will come back to it.
The Architecture in One Paragraph
The cache is ciphertext everywhere except inside the GPU. Blocks are encrypted before they leave GPU memory and stay encrypted through DRAM, across the fabric, and at rest on NVMe. The storage layer, Inferra™, is deliberately blind: it holds encrypted blocks and their integrity metadata, and it holds no keys. It is a place to put ciphertext, and it is architecturally incapable of being anything else.
The consequence is what makes this worth building. Compromise the storage vendor and you hold ciphertext. Compromise the operator, the BMC, a NIC’s firmware, or the fabric, and you hold ciphertext. The blast radius of every component outside the GPU boundary is reduced to encrypted blocks, which is the only honest way to make a retention claim about infrastructure with as many participants as a modern inference path.
This inverts the usual trust posture. Normally the storage vendor asks to be trusted. Here the storage vendor asks to be irrelevant. The security property should not depend on our good behavior, and a design that does is weaker regardless of how trustworthy we are.
Post 13 covers the cryptographic boundary in detail: where exactly the line is drawn, AES-256-XTS at rest and in DRAM bounce buffers, encryption before the DMA on RDMA, and TPM-sealed keys bound to a PCR policy.
Enforcement Belongs on the Client
Here is the unusual part, and the part we would most like argued with.
If the provider enforces the policy, the tenant is still trusting the provider, now about the enforcement rather than about the retention. That is a smaller ask but the same shape, and it does not survive the threat model that motivates ZDR in the first place: the tenant cannot see inside the provider.
So enforcement lives on the client. A ZDR request handler sits in the tool the user actually types into (OpenCode, OpenWebUI, or any OpenAI-API-compatible client) and requires a valid attestation on every turn. Concretely: the client asks for evidence, checks it against policy, and only then transmits the prompt.
A provider that cannot attest never receives the prompt. Not “receives it and is trusted to discard it.” Never receives it. The failure mode is that the request does not happen, which is the only failure mode that actually protects the data.
Two enforcement levels, set by the organization rather than the user:
- Advisory. The client warns and proceeds. Appropriate while providers are still rolling out support, and useful for measuring how much of your traffic could be covered today.
- Mandatory. The client refuses. The prompt does not leave the machine.
Set centrally via MDM, alongside, and conceptually identical to, full-disk-encryption policy. An organization already says “laptops must have FileVault enabled”; this is “agent clients must require attested ZDR,” administered the same way, by the same people, with the same posture-reporting.
That symmetry is intentional. It puts the control in a category that enterprise security teams already know how to operate, rather than inventing a new category that needs its own governance.
What Changes for the Buyer
The change is in the nature of the recourse, and it is worth stating boldly:
| Today | With attested ZDR | |
|---|---|---|
| Enforcement | Contractual | Technical |
| Checked | Annually, by sample | Per request, before disclosure |
| On failure | Breach notification, litigation | Request refused; nothing disclosed |
| Timing | After the fact | Before the fact |
| Evidence | Audit report | Signed attestation per turn |
The current model’s deepest flaw is not that it can be violated. It is that a violation may never be detected. Data disclosed through a misconfigured cache tier leaves no trace for the tenant; the provider may not notice either. The remedy is legal, arrives late, and cannot un-disclose anything.
A technical control shifts the question from what I do after a breach to whether the request proceeded. For a regulated buyer, that is not an incremental improvement in assurance. It is a different kind of statement to put in front of a regulator, and they can check it too.
What is Genuinely Unresolved
Publishing a design at this stage is only useful if you state the open questions as openly as the claims.
Latency. The target is under 10 ms added to time-to-first-token. Attestation is a round trip plus verification, and on a short turn that is a visible fraction of the budget. Caching attestations across turns weakens the per-turn guarantee; not caching them costs latency on every turn. Where that line sits is not settled.
Granularity. Per turn, per session, or per block? Per block is the strongest and the most expensive. Per session is cheap and leaves a window in which platform state could change mid-session. Per turn is the current proposal, and it is a compromise rather than an obviously correct answer.
Key delivery. Getting the sealed key to the GPU boundary without exposing it to the host is the hardest unsolved piece, and it depends on capabilities that differ across accelerator vendors and generations. This is the item most likely to constrain what ships first.
Revocation and rotation. What happens when a provider’s platform measurement changes because they patched something? Too strict and every legitimate security update breaks every client. Too loose and the measurement stops meaning anything.
Ecosystem. A client-side control is only as useful as the number of clients that implement it and providers that can satisfy it. This needs to be a specification with multiple independent implementations, not one vendor’s feature.
What We Are Asking For
Not a purchase. A co-author.
The specification is more valuable with security engineers, inference providers, and enterprise buyers arguing about it than it is with us writing it alone, particularly on the five open questions above, where the right answers depend on operational realities we do not all share.
If you sign ZDR clauses, or you are a provider who would have to satisfy this, or you maintain a client that would implement the handler: the reference work is in the High Assurance Agent Init repository, and the specification is open for comment.
The immediate practical step, regardless of whether any of this ships: go and ask your current provider the three questions from post 4. Where does my KV cache live and for how long; is it encrypted at rest and on the wire, and who holds the keys; and what evidence can you give me, per request, that both are held? The third question is what this post exists to answer.