Cartridges: A Cache You Train, Not a Context You Fill

Источник: Everpure Blog

Cartridges: A Cache You Train, Not a Context You Fill

Source: Everpure Blog

Cartridges: A Cache You Train, Not a Context You Fill by Everpure Blog This article explores what cartridges are, how they use compact KV caches to streamline long-context inference, and how they can help reduce resource demands. The post Cartridges: A Cache You Train, Not a Context You Fill…

•Updated: October 2, 2026

Imagine an AI agent embedded in an investment firm, keeping track of every company the firm has money in. Behind it sits years of data on each one: financial filings, calls with company leadership, and notes from past meetings.

Someone on the team asks how one company’s margins broke down by segment last quarter. To answer, the agent reads through the company’s filing. An hour later, someone else asks about that same company’s headcount, and the agent reads through the same filing again. A third person asks who that company’s competitors are, and the agent reads it again. The filing never changed, yet the model started over from scratch every time.

Now a harder question comes in: Which companies in the portfolio have mentioned pulling back on hiring this year? This isn’t a rereading problem anymore. Answering it well means holding many companies’ filings in the model’s memory at once, not just one, and there’s a limit on how much a model can hold in a single conversation before it runs out of space.

This is the shape of most real long-context work. Some of it is one document, asked about again and again. Some of it is many documents, needed all at once. Either way, what varies between requests is a few hundred tokens of a question, but what stays constant is millions of tokens of data behind it. Yet the dominant way we give models context treats all of that data as new every time, whether it’s one filing or 20.

The cost of that shows up in the KV cache. Caching keys and values is what makes autoregressive decoding fast, but the cache grows linearly with input length and lives in GPU memory for the duration of the request. Answering one question about a hundred-thousand-token filing isn’t expensive because the model is thinking hard. It’s expensive because it’s holding all hundred-thousand tokens of cache in memory while it does.

At Everpure, we spend a lot of time on this problem from the infrastructure side: where KV cache lives, how it moves in and out of GPU memory, and how to avoid recomputing it. Cartridges are a research approach that comes at the same problem from a different direction. They aim to keep the benefit of reading the filings without the cost of keeping them in memory.

What we do today

Every enterprise sitting on a pile of data runs into some version of this problem from the outset. Sometimes it’s a filing, and sometimes it’s a codebase, support ticket history, or years of case files. The data itself barely changes, but what people ask about it changes constantly. Regardless, the model needs the relevant data in front of it somehow, and there are three dominant ways to get it there today, each with pros and cons.

The most direct approach is in-context learning. Put the documents straight into the model’s context window and let it answer from there. This generally produces the highest-quality responses of the three, because nothing about the documents gets lost or approximated on the way in. The tradeoff is on the hardware side. Every question pays for prefill compute and cache memory again, and that cache has to sit in the GPU’s High Bandwidth Memory (HBM), the same finite memory the model’s weights already need. Once it fills up, the only way to keep going is to summarize what’s already happened, trading away the high quality that made this approach worth choosing in the first place.

The second method is retrieval. Instead of handing over an entire document, this method fetches the handful of paragraphs that look most relevant to the question and hands over just those. It’s cheap, mature technology that handles document changes gracefully. The weakness is that retrieval only works when a question’s answer lives inside one identifiable chunk of text. “Which companies in the portfolio have mentioned pulling back on hiring this year?” fails that test, because no single paragraph contains the answer. Answering it means checking every company’s data, not just the one paragraph that best matches the question. That isn’t what retrieval was designed for.

The third way is fine-tuning. This method bakes the document’s information directly into the model’s weights, so nothing needs to be handed over at inference time. Once trained, this is the only one of the three methods that adds no per-request cache memory, since the knowledge lives in the weights instead of a cache. Getting there is a different matter. It means building a data pipeline and running real training, with the risk that training hard on one company’s documents degrades what the model already knew about everything else. It also only works for information that holds still. The moment a document changes, the model’s knowledge of it is out of date, and the only fix is training it again.

Cartridges may soon prove to be a fourth approach. They don’t fetch fragments the way retrieval does, and they don’t touch the model’s own weights the way fine-tuning does. What they aim to keep is in-context learning’s high-quality responses and the accuracy that comes from having the whole document available. What they replace is the expensive part of loading the KV cache of each document during every agentic conversation. A cartridge trains a much smaller, highly specialized cache just once, offline, and reuses it from then on, saving HBM.

What a cartridge is

A cartridge is that smaller cache we just talked about. Instead of computing keys and values by reading the document when needed, you train a fixed number of key-value pairs directly, offline, but far smaller in size than the cache of the document itself. Those trained pairs get loaded straight into the model’s attention layers, in place of whatever a full read of the document would have produced. A cartridge’s whole purpose is being small, so it can be roughly ~40x smaller than the cache of the document it replaces, depending on how it’s configured. And once it’s trained, it’s still just a KV cache, not some new kind of object. To an inference server, a cartridge looks like any other KV cache, so in principle, it can be served without changes to the serving stack. That also means infrastructure built to store, move, and reuse KV caches is relevant to cartridges.

Training a cartridge sounds like it should be simple: just have it predict the document’s next word, the same way a model is normally trained. That falls far short of in-context learning’s quality. It’s the same failure as a student who’s memorized an equation well enough to write it out from memory but doesn’t know when to use it or what problem it actually solves. When trained this way, the model has the document’s surface text down cold. It just can’t use any of it to answer a question about it.

What works instead is a process called self-study. First, break the document into chunks. Then, for each chunk, have the model quiz itself, generating a synthetic conversation about that piece of the document. Sometimes the conversation is about a question, a summary, a comparison, or specific facts pulled out and organized. The more this simulates real-world conversations, the better. Finally, train the cartridge on all of those conversations using something called context-distillation. Instead of training toward repeating the document’s actual text, you train toward the answer a model would give if that chunk were sitting right in its prompt, the same way in-context learning would.

That’s the whole idea. You’re not teaching the model the document; you’re teaching it to act as though it had just read the document. This is why the quizzing needs real variety, because a cartridge only gets good at the kinds of questions it was actually trained on. Training it on a wide mix is what lets it hold up against real questions later.

In research from Stanford’s Hazy Research, cartridges trained with self-study matched in-context learning quality on challenging long-context benchmarks while using 38.6X less memory and enabling 26.4X higher throughput. In practice, this means cartridges could let the same hardware serve far more of these conversations at once, because each one only needs a small cache instead of a full-size one. Results in production will depend on the model, workload, and configuration.

Not every document is worth training a cartridge for

Training a cartridge isn’t free. It costs real compute, and that cost is paid before anyone has asked a single question.

That’s fine, because the cost gets paid back over every question that follows. Without a cartridge, the expensive part is holding a full-size cache in memory for every single question that comes in. Self-study trades that recurring cost for a single upfront one: Pay once to train the cartridge, and every question after that works from a much smaller cache. If a document only gets read once, that’s a bad trade to make. If a document keeps getting reused by team members asking it new questions for the next year and a half, it’s an easy one.

The real test for any document is how often your team will actually ask AI questions about it and how likely the document is to change while they’re still asking. A document your team keeps querying, one that stays put, is exactly the case cartridges are built for. A document someone asks about once, or one that’s modified every week, is not.

There’s already early evidence this trade pays off outside a research paper. Engram, working with the legal-AI company Harvey, built an agent that studies a synthetic law firm’s entire set of files up front, then completes long tasks based on them at close to a tenth of the cost of a frontier model, with a comparable or better success rate. Their method costs $0.13 per query at a 30% success rate compared to $1.32 and 25% for Claude Opus 4.8, based on OpenRouter pricing as of August 16, 2026. Although Engram’s mechanism differs from cartridges (it learns into model weights and text notes rather than a trained cache), it cites the cartridges research as an influence. The bet is the same one: absorb the firm’s documents once, offline, so that queries afterward don’t have to start from scratch.

One notable result from Engram’s work is a closed-book test. Before any training, they asked the base model questions about the firm with nothing to search and no notes to check; it answered correctly 4.7% of the time. After training, with the same constraints, accuracy climbed to 72.6%. Nothing about the test changed. What changed is that the agent had absorbed the firm’s knowledge ahead of time. It’s the same principle behind cartridges: do the learning once, ahead of time, rather than rebuilding context from scratch on every request.

Where cartridges differ from an ordinary KV cache

Some questions can’t be answered from a single document alone. Did what a company’s leadership said on the earnings call actually match what showed up in the quarter’s filing? Answering that means having both documents in view at the same time.

The original cartridges research found something surprising here. Train one cartridge on the filing, train a separate cartridge on the call notes, entirely independently of each other, and you can place both in front of the model at inference time without retraining either one.

That’s harder to do with ordinary KV caches. Researchers studying this exact situation, combining several separately cached documents at once, found that precomputing each document’s cache on its own, then just pasting them together, can degrade answer quality. Each cache was built as if it were the only text the model would ever see, so the filing and the call notes never learned anything about each other, and pasting them together doesn’t fix that gap.

What’s still unsolved

Cartridges aren’t ready to drop into any setup today for several reasons.

The first real hurdle is that a cartridge is trained against one specific model’s attention layers, using real, dedicated GPU compute, and none of that work transfers between models. Swap in a different model, even a newer version of the same one, and the old cartridge stops working, so a fresh one has to be trained from scratch. This is also why cartridges only work with a model you run yourself; there’s no way to inject one into a model reached only through a hosted API. That access barrier matters less as more enterprises explore running open-weight models themselves. But that shift doesn’t solve the model transferability problem. Most enterprises run more than one model, for different tasks or different vendors, and every one of those models needs its own complete set of cartridges, built on dedicated hardware, and kept in sync every time a model gets upgraded. That’s a real ongoing cost, and one that has to be weighed against the savings. It’s also a data management problem as much as a training one: Someone has to track, version, store, and serve the right cartridge for the right model. That’s the layer of the stack where Everpure focuses.

There’s also nothing here that handles a document that changes. Every result so far assumes the material is static. The moment a filing gets amended, the cartridge built from it is out of date, and the only fix is training a new one. For now, cartridges only make sense for documents that are genuinely static, not ones getting revised regularly. This is one area where retrieval currently has a clear advantage.

The tooling isn’t ready for enterprise scale either. This is still a research result, not an engineering one, so building a cartridge today takes highly specialized AI engineers, the kind of dedicated team most enterprises don’t have sitting around free to take this on. That gap between a research result and something a team can actually use is exactly what we’re working on closing.

None of this changes the underlying case for cartridges: The tradeoffs are real, but so is the upside, and that’s reason enough to keep pushing on them. We’re doing that work ourselves here at Everpure. We’re actively exploring how cartridges fit into enterprise AI infrastructure and what it takes to make them usable at scale. We’ll share more as that work matures. Stay tuned.

Cartridges are still an emerging research area, but the runtime side of the problem is one you can address today. To see how Everpure helps store and reuse KV caches across requests instead of recomputing them, read our blogs “Architecting for Reuse: A Deep Journey into the Heart of KV Caching” and “Pure KVA 1.0 Is GA: Inside the Architecture of a Fully Managed Enterprise KV Cache Solution.”

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.