The Architecture Behind More Efficient AI Inference

Source: Everpure Blog•

The Architecture Behind More Efficient AI Inference

The Architecture Behind More Efficient AI Inference by Everpure Blog Explore how shared KV caching, intelligent request routing, and hybrid model deployment can improve on-premises AI inference efficiency and reduce effective cost per token. The post The Architecture Behind More Efficient AI…

Over the past year, the economics of inference have pushed more enterprises to look seriously at on-premises capacity. It’s the challenge behind the AI token optimization reference architecture we announced at Everpure Accelerate London last week. Frontier output prices had jumped somewhere between two and seven times over, GPUs were scarce, and inference demand was doubling every few months. So enterprises did the sensible thing and started pulling workloads onto hardware they owned. The logic was straightforward: buy the GPUs, install an inference engine, run your own models, and stop paying for every token at a markup.

The catch is that on-prem costs are harder to see. In the cloud, a bigger invoice tells finance something changed. On-prem, there’s no invoice to read. The first signal usually comes from engineers, when responses feel slow or the system seems to need more hardware than it should. Someone finally opens the dashboard and sees GPU utilization pinned near 100%, which looks like exactly what the capacity plan was supposed to deliver.

But utilization only says the chips are busy. It says nothing about what each answer costs. For that, you need a different number: the effective price per token, meaning everything that went into producing your tokens (GPU hours, power, the rest) divided by the tokens you got out. More teams are starting to run that math, and many find that with the wrong inference stack, a “free” open-weight model can cost as much per token as a frontier API.

High utilization and low efficiency can sit on the same cluster at the same time. That gap is where most of the on-prem inference story gets missed, and closing it means looking past the model to the rest of the stack around it.

High GPU utilization is not necessarily good GPU utilization

Picture the GPU fleet as a restaurant kitchen during the dinner rush. Every burner is lit; every station is moving. From the dining room, it looks like a kitchen at peak performance. Walk inside, and you find something stranger: Nothing in this kitchen gets prepped ahead and reused. Every stock gets simmered from raw bones and every sauce gets built from scratch, for every single ticket, even when the dish that just left used the exact same stock. The kitchen looks busy, but it isn’t productive. Most of that motion is redoing work that already got finished five minutes ago, one station over.

That’s what high GPU utilization tends to hide. There’s a comforting assumption that a GPU running hot is a GPU earning its keep. It isn’t, necessarily. A GPU generating new tokens is doing the work you bought it for. A GPU re-deriving a stock that another node already made is cooking from raw ingredients when the kitchen already had what it needed on the stove.

Where the GPU time actually goes

To see why that repeated prep costs so much, it helps to know where a model actually spends its time. Before it produces a single token of output, it has to work through the entire prompt first, a step called prefill. Prefill scales quadratically with length: double the prompt and you roughly quadruple the work. A 100,000-token prompt on a large model, even running on current hardware such as NVIDIA DGX B300™ systems or NVIDIA Vera Rubin NVL72 systems, can still burn 8 to 10 seconds of GPU time before the first output token exists.

That problem gets much worse with agentic AI. Five hundred engineers may each launch coding agents against the same codebase, using the same system prompts, tool definitions, and large chunks of shared context. The relevant context is nearly identical every time, but a naive setup rebuilds it for each request, spending 4,000 to 5,000 GPU-seconds recomputing the same KV state. A round-robin load balancer makes it worse rather than better, because the next request from the same session lands on a different node with a cold cache and pays the full prep cost again.

This is the part of token economics that the headline number misses. Frontier APIs price a cache hit directly: Reused context costs a fraction of a fresh input token, so heavy reuse shows up as a real discount on the invoice. On-prem, there’s no separate line item for a cache miss, because there’s no invoice at all. The cost shows up as lost throughput instead. Every GPU-second spent re-deriving context someone already computed is a GPU-second not spent generating new tokens. That quietly raises the cost of every token you do produce, and tokens-per-dollar goes down, even though nothing on a dashboard flags it as a cache problem.

Compute once, serve everywhere

Think of it as the system architecture equivalent of mise en place: prep the components once at the start, and let every station pull from the same tray. In inference terms, the key-value (KV) cache generated by reading a prompt shouldn’t be trapped on the single GPU node that built it. It should be computed once and made available to the entire fleet.

By moving from a per-node cache to a fleet-wide shared cache, any node can leverage existing work. NVIDIA Dynamo and Pure Key-Value Accelerator (KVA) are the pieces of infrastructure that make that handoff possible, moving KV state between nodes instead of pinning it to whichever GPU happened to build it first. When a prefix has already been computed somewhere in the cluster, the system can reuse that KV cache from the appropriate cache tier instead of recomputing it from scratch. In Pure KVA benchmark testing, a roughly 10-second recomputation can become an approximately 500-millisecond cache injection, while freeing GPU cycles for new token generation.

The same architecture also addresses another source of waste: idle sessions. Agentic workflows spend much of their time waiting on tool calls, build processes, or external API responses. Holding expensive GPU memory for a paused session is wasteful. Instead, the system can demote an idle KV cache to shared storage and restore it without recomputation on whichever node has capacity when work resumes. All of this happens transparently behind a single gateway, requiring zero changes to client code. Avoiding recomputation is one half of the efficiency equation. The other is making sure every request runs on the right capacity in the first place.

On-prem and frontier, not on-prem or frontier

Reusing computation only addresses part of the economics. An ideal platform also needs to dynamically decide where each request should run. While a naive on-premises design lowers the cost of bulk workloads, it traps those workloads on fixed hardware, locking them out of frontier capabilities when a complex task demands it.

Breaking that constraint requires unifying the two environments.

By putting a single gateway in front of both on-premises and external models, the architecture abstracts away the hardware entirely. Before a request goes anywhere, it gets sized by a lightweight difficulty estimator that looks at what the prompt is actually asking for and routes it to the model that can handle it. This pattern is showing up industry-wide, not just here. NVIDIA’s NeMo Switchyard, an open source routing layer built on the same idea, has already been benchmarked by LangChain at a 74% cost reduction by routing between an efficient open model and a frontier one, sending only a small fraction of calls to the frontier tier.

NVIDIA’s leadership has reached a similar conclusion. Jensen Huang has been candid that self-hosting carries real costs of its own: the expertise, guardrails, and maintenance that a closed API requires. He also tells his own employees to rent cloud intelligence wherever that’s the faster path. The gateway we described above is what lets an enterprise act on that, routing the bulk of traffic to owned infrastructure while still reaching for a frontier model exactly when the task or the moment calls for it.

The same routing also covers timing, not just difficulty. When traffic spikes past what the owned fleet can absorb without slowing everyone else down, the gateway can send the overflow to a frontier model instead of letting queue depth build on hardware that’s already maxed out. Capacity you own stays sized for steady-state load. Capacity that you rent covers the peaks. The users still see one system throughout, while the platform reserves paid frontier capacity for exactly the load and exactly the difficulty that calls for it.

But won’t cheaper models make this moot?

A reasonable objection is that token prices will keep falling and context windows will keep growing. Doesn’t that make recomputation waste less important over time?

In practice, adoption makes the problem larger, not smaller. Cheaper and faster are not the same claim, though the objection treats them as one. A model getting cheaper per token or a context window large enough to hold an entire codebase says nothing about how quickly that model produces the tokens you’re paying for. Prefill still scales quadratically with prompt length, so a bigger context window makes that wall taller, not shorter, every time a request reads through it from the beginning. What actually buys you speed, meaning more tokens generated per second on the same GPU, is reusing a KV cache that already exists somewhere in the fleet instead of recomputing it. And that speed is what turns into a better tokens-per-dollar number, the same metric from the opening of this post, regardless of what a model’s sticker price does over the next year. More throughput per GPU also means more users served on the fleet you already own, which is the number that was supposed to improve when you moved on-prem in the first place.

Add capacity last

When inference costs stay high, the instinct is often to add capacity: faster GPUs, more nodes, a bigger cluster. That instinct skips a step. Efficient inference is not one trick; it’s an entire stack. It starts with knowing how much intelligence a given prompt actually needs and routing it to the model sized for that. It continues with a KV cache available to the whole fleet. And it ends with a release valve: the requests that genuinely need frontier-level capability or arrive during a spike in usage your fleet cannot absorb go to a frontier model on demand. Put those three together, and you get a system that’s actually earning the word efficient. High utilization alone doesn’t tell you whether that’s what you’re looking at or whether it’s just a fleet working harder unnecessarily.

Get all three of those right first, and capacity becomes a much smaller, much more precise decision. The same gateway that routes hard tasks to a frontier model on demand also means peak load doesn’t have to be absorbed by owned hardware sitting idle the rest of the week. That turns growing the GPU footprint into something you do deliberately, once routing and caching have already told you where the real bottleneck is. Buy the additional GPUs after the stack is coordinated, and every one of them buys real capacity instead of quietly covering for unnecessary waste.

This is the same bet the industry is now making about the right shape for AI at scale: reserve frontier-scale capability for the problems that actually need it, and run everything else on efficient, specialized models close to where the work happens.

At Everpure Accelerate London last week, we announced the AI token optimization reference architecture to put these ideas into practice. To see how to implement fleet-wide cache sharing and hybrid routing in your environment, view the full AI Token Optimization for On-Prem Inference reference architecture.

What this article says