How we would like to work with accelerator and memory-fabric vendors
This is an open invitation, written in public on purpose.
The KV cache path, the route a block takes from wherever it is stored to the attention kernel that needs it, is being defined right now, across inference engines, routers, DPUs, and the emerging generation of Ethernet-attached context flash. Most of it is being defined well. But two interfaces don’t exist anywhere in the stack, and they determine whether long-context inference is efficient or merely possible.
We would like to co-design both, and we would like to say publicly what we would bring and what we would ask for, because a collaboration proposed in the open is easier to join than one negotiated privately.
Where the pieces already sit
The path today has three reasonably well-defined layers.
The open ecosystem. Routers and orchestrators that decide global locality and intent. Inference engines (vLLM, TensorRT-LLM, SGLang) that own HBM, the block table, and scheduling. And a KV connector boundary where an external cache can attach. This layer is healthy, standardizing, and where the Open KV Cache API belongs: a vendor-neutral contract for block identity, typed metadata, an asynchronous ticketed data plane, and mandatory tenant and session scoping.
Proprietary userspace. Between the connector boundary and the transport sits the part that decides which blocks to move, when, and at what budget. This is where Inferra‘s™ planner runs, and where similar components run for others. It is legitimately competitive: prediction quality is the differentiator, and it should be.
Vendor data planes. Below that are the transport and media: DRAM tiering and cross-node movement, asynchronous byte movement, the storage framework on the target side, and the flash itself.
Everything in that stack exists or is in active development. Nothing described below requires a new layer.
Missing Interface 1: a Prefetch-Aware Extension Point on the Target Side
Today, a capacity tier is exactly that: capacity. Blocks live on flash; the host asks for one; the block comes back. The storage side participates in movement, not placement.
That is a missed opportunity, because the target knows things the host does not. It sees the whole population of blocks, their access history across sessions and nodes, and the physical layout. What it cannot do is act on any of it, because the interface offers no way to express intent beyond “read this block.”
What is missing is a small extension surface that lets a partner attach:
- Locality signatures: a way to say that a set of blocks belongs together and is likely to be needed together, without the target needing to understand what a KV block means semantically.
- Priority and deadline: a request carrying “this is needed by roughly time T” rather than “fetch this now.” Deadline-aware I/O is old technology and it is exactly what latency-hiding needs.
- Statistical prestage into the near tier: the ability for the target to hydrate the DRAM tier from flash on its own statistical grounds, ahead of the request, so the host-side planner is choosing among warm candidates rather than cold ones.
That last one is the substantive ask. A capacity tier that can pre-stage selectively stops being storage and becomes part of the memory hierarchy. It is the difference between a tier that answers questions and a tier that anticipates them, and on the G3.5-class Ethernet-attached context flash platforms now reaching the market, the target has both the compute and the proximity to do it well.
Note what this does not require: no knowledge of model internals on the storage side, no exposure of anyone’s prediction algorithm, no semantic understanding of attention. It is an extension point for hints and deadlines, and each party keeps what it does best.
Missing Interface 2: a Device-Side Attention Signal
The second gap is more fundamental and harder.
Prediction today runs in userspace, from what userspace can see: request context, block metadata, engine timing, observed reuse. That is enough to produce large long-context wins; it is what the current measurements are built on. It is not the whole signal.
The state that actually determines which blocks matter is produced inside the attention computation. Which blocks a sparse or top-k pattern will select is a function of the query against the keys, computed on the accelerator, microseconds before the blocks are needed. By the time that information surfaces to userspace, the window in which prefetching it would have helped has substantially closed.
So there is a structural mismatch: the prediction runs far from the data that produces it. Everything a userspace planner does is an inference about a signal that exists, in exact form, a few hundred microseconds away and one privilege boundary below.
What would close it is some capability-scoped, versioned way for a partner component to observe attention state at source, sampled, aggregated, and exported through a defined interface, so that prediction can run near where the signal is generated rather than downstream of it.
To be precise about what is and is not being asked for: not raw register access, and not a second memory manager. A narrow, versioned, capability-scoped service interface. The accelerator vendor owns the silicon, the firmware, and the signing; what is being requested is a defined surface, not a way around one. Ownership stays exactly where it already is.
What we would bring
An open contract, kept genuinely open. The Open KV Cache API is vendor-neutral, usable without any particular vendor’s hardware, and governed in the open. We carry the community work, recruit storage partners, and maintain conformance tooling. We are not asking any engine to fork, and we are not asking anyone to adopt a proprietary interface.
A reference implementation. Inferra implements that contract end to end (target integration, planner, and connector) and we intend for it to stay ahead of other implementations of the same contract by being better, not by being the only one.
Reproducible benchmarks and observability, published jointly. Measured at the token and query level, not in storage microbenchmarks. A prefetch mechanism that improves a storage benchmark and not time-to-first-token has not improved anything that matters.
The security work. The same primitives that make this fast make it auditable: deterministic block identity gives integrity by construction, mandatory tenant scoping makes cross-tenant reuse impossible by construction rather than by policy, and a defined boundary is the only place encryption and attestation can attach at all. That matters increasingly to public-sector and regulated buyers, and it is not achievable inside a hand-rolled copy loop.
How we would want to run it
Four principles, offered as constraints on ourselves as much as on anyone else.
Open at the contract layer, competitive above and below it. The interface should be neutral and standardizable. Prediction quality and device implementation are where anyone should be free to compete. A specification that requires a particular vendor’s hardware is not a specification.
Evidence before each deeper step. Each stage should ship something useful on its own and carry an explicit exit gate: measurable long-context gain with no short-context regression; sampling overhead below the latency it removes; a device-side path whose advantage is real rather than assumed. Nothing deeper begins until the shallower step has produced evidence, and no software capability should be gated on firmware access.
Ordinary prefill stays the correctness fallback, always. Every optimization here is a performance path with a correct, boring alternative underneath it. If prediction is wrong, the system computes the answer the slow way and is still right. Never trade away that property; it is what makes it safe to be aggressive above it.
Publish the methodology with the numbers, and always with the regime. A multiple without a sequence length is not a result. These gains scale with context because prefill is superlinear in sequence length while restore is linear, so a joint benchmark should state the context lengths tested, the reuse share, and the effect on the whole turn as well as on time-to-first-token. At multi-million-token contexts, prefill is nearly the entire turn and those two converge; at short contexts, they do not. Saying which regime a number came from is what makes the long-context result credible.
Why now
These seams are being cut now. Router intent is being normalized into typed hint protocols. KV connector interfaces are hardening across engines. Ethernet-attached context flash is arriving as a product category. Within a couple of release cycles these interfaces will be set, and the predictive layer between them will either have an owner and a specification, or it will be twenty incompatible internal implementations that no security control can attach to and no second vendor can implement.
The second outcome is bad for everyone, including the vendors whose hardware would be underused.
If you work on accelerators, DPUs, memory fabrics, context-class flash, or an inference engine, particularly if you think one of the two interfaces above is wrong, or is the wrong shape, that is the most useful conversation available. The specification is open for comment, the reference implementation is real, and the benchmark methodology is meant to be argued with before agreement.
Part of an ongoing series. See Follow-up on Paged Attention over RDMA for why prediction, rather than transport, turned out to be the binding constraint, and Why Inference Needs a Rethink for the argument that the KV cache is a storage problem wearing a GPU costume.