MiniMax M3 Prompt Caching on SambaCloud: How It Works

Source: SambaNova•

MiniMax M3 Prompt Caching on SambaCloud: How It Works

Prompt caching, first introduced on SambaCloud with MiniMax M2.7 , is now available for MiniMax M3. When requests share a stable prefix of at least 4,096 tokens, SambaCloud can serve that prefix from cache instead of recomputing it, with no code changes required. Cached tokens are billed at 90%…

Prompt caching, first introduced on SambaCloud with MiniMax M2.7, is now available for MiniMax M3. When requests share a stable prefix of at least 4,096 tokens, SambaCloud can serve that prefix from cache instead of recomputing it, with no code changes required. Cached tokens are billed at 90% below the standard input rate, and across 8k to 192k tokens of context, cache hits cut time to first token (TTFT) by 35% to 88%.

TL;DR

  • Prompt caching is now available for MiniMax M3 on SambaCloud, so when agents resend the same long context, such as a repository, a set of research papers, or a fixed set of tool definitions, and the prefix is served from cache, only the new tokens are processed fresh.

Prompt caching is now available for MiniMax M3 on SambaCloud, so when agents resend the same long context, such as a repository, a set of research papers, or a fixed set of tool definitions, and the prefix is served from cache, only the new tokens are processed fresh.

  • It runs on Automatic Prefix Caching and needs no setup: prefixes of at least 4,096 tokens qualify, up to a maximum cacheable length of 192,000 tokens, and under sustained traffic hit rates typically climb above 90%.

It runs on Automatic Prefix Caching and needs no setup: prefixes of at least 4,096 tokens qualify, up to a maximum cacheable length of 192,000 tokens, and under sustained traffic hit rates typically climb above 90%.

  • Across 8k to 192k tokens of context, cache hits cut TTFT by 35% at the short end and 88% at the long end, a median speedup of 4.7x, with TTFT dropping from 8.4 seconds to 1.0 seconds at 192k tokens.

Across 8k to 192k tokens of context, cache hits cut TTFT by 35% at the short end and 88% at the long end, a median speedup of 4.7x, with TTFT dropping from 8.4 seconds to 1.0 seconds at 192k tokens.

  • Cached tokens on MiniMax M3 are billed 90% below the standard input rate, at $0.06 per million tokens instead of $0.60.

Cached tokens on MiniMax M3 are billed 90% below the standard input rate, at $0.06 per million tokens instead of $0.60.

  • Every response returns a prompt_tokens_details object with cached_tokens and cache_creation_tokens, so you can confirm exactly how many tokens were served from cache.

Every response returns a prompt_tokens_details object with cached_tokens and cache_creation_tokens, so you can confirm exactly how many tokens were served from cache.

Why Prompt Caching Matters More with MiniMax M3

MiniMax M3 is built for long-horizon agents. Its 1M-token context window, powered by MiniMax Sparse Attention, means whole repositories, multi-day agent logs, and full research papers fit in a single request. Long, repeated contexts like these are where prompt caching pays off. Caching covers prefixes of up to 192,000 tokens, and within that limit, the longer the stable prefix, the larger the share of each request that can be served from cache instead of recomputed.

Consider what an M3 agent loop looks like in production:

  • A coding agent that loads a repository once and then issues dozens or hundreds of edit-and-test turns against it.
  • A research assistant that holds a set of papers in context and answers a stream of questions about them.
  • A computer-use agent that carries the same operating instructions and tool definitions through every step of a long task.

Without caching, each of those turns reprocesses the same context from scratch. With caching, only the new tokens, the latest user message or the newest tool result, are processed fresh. In production, agentic workloads typically see cache hit rates around 90% on average, per industry estimates.

Prompt caching on SambaCloud is powered by Automatic Prefix Caching. When a new request's leading tokens match a prefix we've recently processed, we serve those tokens from cache rather than recomputing them. Nothing is required on your end beyond keeping the prefix stable.

To maximize cache hits:

  • Put stable content first. System prompt, documents, examples, and tool definitions go at the top; the parts that change go at the end.
  • Keep the prefix byte-identical. Rewording the system prompt, inserting a message before it, or reordering content starts a new cache entry. Changing only the user message reuses the cached prefix.

The first few requests populate the cache; savings compound as the prefix is seen repeatedly. For sustained traffic, hit rates typically climb above 90%.

What Prompt Caching Means for Latency

Across 8k to 192k tokens of context, cache hits cut time to first token (TTFT) by 35% at the short end and 88% at the long end, a median speedup of 4.7x. At 192k tokens, TTFT drops from 8.4 seconds to 1.0 seconds.

What It Means for your Bill

Cached tokens on MiniMax M3 are billed at 90% below the standard input rate.

See It in the Response

Every response includes a prompt_tokens_details object:

cached_tokens is the count served from cache and billed at the discounted rate. cache_creation_tokens is non-zero on the request that writes the cache entry and zero on subsequent hits; it is informational and carries no extra charge. prompt_tokens − cached_tokens is what was processed fresh.

Good to Know

  • Prompt caching is available on MiniMax M3. Other models return cached_tokens: 0.
  • Cache state is local to each serving node; on multi-instance deployments a prefix is cached independently per node.
  • Eviction is LRU and depends on load. Cache persistence is not guaranteed, so the time to live for a given prefix varies with traffic.

Get Started with Prompt Caching on MiniMax M3

If your workload sends the same long context repeatedly, caching is already working for you on M3. Check prompt_tokens_details in your next response to confirm. Full details are in the prompt caching documentation.

New to SambaCloud? Explore MiniMax M3 in the playground, or generate an API key and start building.

What this article says