Kimi K3 on AMD Instinct GPUs with TokenSpeed

Source: AMD•

Kimi K3 on AMD Instinct GPUs with TokenSpeed

TokenSpeed delivers fast, agentic LLM inference on AMD Instinct™ MI355X GPUs, with optimized kernels and leading Kimi K3 performance.

TL;DR

  • Multi-vendor agentic inference, with AMD at the core: TokenSpeed combines a C++ scheduler, Python execution plane, and modular kernel APIs, with unified cache management and execution across accelerators.
  • Leading Kimi K3 performance on 8x AMD Instinct™ MI355X GPU: 216 output tokens/s per user at concurrency 1, 1.4x of ATOM, on par performance at higher concurrency, with higher precision.
  • Leveraging Triton foundation, agent-assisted Gluon optimization: Specialized kernels, fusion, and Iris communication deliver up to 5.1x faster MLA prefill, 2.1× faster speculative attention, and 4.4× faster MoE all-reduce against the baselines detailed below.

TokenSpeed: Agentic Inference with AMD at its core

TokenSpeed is a recent open-source LLM inference engine designed for low-latency generation on agentic workloads. Coding agents repeatedly return to long conversations, process new tool results, and generate the next action while the user waits. Serving them well requires low decode latency, efficient prefix reuse, and predictable scheduling.

TokenSpeed gives each part of that problem a clear home. A C++ control plane manages request lifecycles and cache ownership as a finite-state machine, with the type system enforcing safe resource reuse. A Python execution plane supports rapid model development. A layered kernel subsystem connects portable model code to specialized implementations through shared APIs, a central registry, and an accelerator plugin mechanism. Prefix caching, speculative decoding, graph execution, and disaggregated serving build on these common mechanisms.

AMD support is fundamental to this design. TokenSpeed enabled Kimi K3 on CDNA4 at Day 0, with prefix caching, speculation, disaggregated serving, and graph-captured decode. TML Inkling also arrived at Day 0 with AMD support, an AMD Quark MXFP4 checkpoint, and Gluon specialized kernels. The same kernel system now covers MHA, MLA, DSA, KDA, MoE, and GEMM on AMD. These results show how common model and scheduling infrastructure can support new architectures while specialized kernels exploit the hardware.

Kimi K3 is a frontier agentic model, and this post shows how TokenSpeed serves its long conversations on MI355X, especially for high interactivity. The multi-turn measurements below report 216 output tokens/s per user at concurrency 1 on eight MI355X GPUs.

Our optimization process starts with portable Triton implementations as a functional baseline and numerical reference. Profiling then guides agent-assisted development of specialized Gluon kernels that implement the same operator interfaces and meet the same numerical contracts. The approach extends across the stack: combining projections, retaining intermediates on chip, and fusing computation into Iris kernels, a Triton-extension that enables multi-GPU kernel programming.

Kimi K3 Performance on Long Agentic Conversations

Agentic coding workloads repeatedly revisit a long conversation while the user waits for the next tool call. To measure this pattern, we replayed multi-turn SWE-smith coding sessions with TokenSpeed’s open-source agentic benchmark. Each conversation opens with a ~50K-token prompt, adds ~800 tokens per turn over 10–15 turns, and generates 500 tokens per reply. The existing run uses eight AMD Instinct MI355X GPUs, TP8, EAGLE3, FP8 KV, and an approximately 90% cache-hit rate.

With one active session, this run reports 216 output tokens/s per user, or 4.62 ms per output token (TPOT), and a 0.62-second average time to first token (TTFT). At this high interactivity situation, which the industry is heavily pushing towards, it achieved 1.4× compared to ATOM. Performance is on par with or slightly faster than ATOM as we scale up to concurrency 2, 4, and 8. At concurrency 16 (C16), each agent receives 39 output tokens/s.

Figure 1: Kimi K3 agentic serving on eight MI355X GPUs. Per-user decode speed is 1000 divided by time per output token in milliseconds; it excludes time to first token and queuing. Total-token throughput counts input and output tokens, including cached prompts. Each point is a single run at one concurrency setting; serving conditions are described in the configuration section.

ATOM leads at C16. At C16 ATOM switches to a different scheduling approach termed prefill coalescing. This delays prefill requests, allowing more decode requests through, until enough prefill requests are available to be batched. This trades TTFT (1.44 -> 2.90s mean) for TPOT (25.17 -> 19.28ms mean) compared to default split scheduling and results in an overall throughput win, however p95 user request TTFT grows 3.54x (2.95 -> 10.46s) in our measurements. ATOM also conducts more aggressive quantization as shown in the following table:

Model Component

TokenSpeed

ATOM

Attention Q/K/V and output projections

BF16/BF16

FP8/FP8

Shared-expert projections

BF16/BF16

FP8/FP8

Routed-expert GEMMs

FP8/MXFP4 (small decode); MXFP4/MXFP4 (prefill)

MXFP4/MXFP4 (prefill)

KDA recurrent state

FP32

FP16

MLA KV cache

FP8

FP8

Kimi K3 Architecture and the Eight GPU Execution Path

Large language models generate text one token at a time. Each new token flows through attention, which reads the conversation, and a feed-forward network, which transforms the current representation. Kimi K3 has 93 decoder layers: 69 use Kimi Delta Attention (KDA), and 24 use gated Multi-head Latent Attention (MLA). Its first feed-forward layer is dense; the remaining 92 use Latent MoE. Attention Residuals (AttnRes) connect representations across layer blocks (Figure 2).

Figure 2: Kimi K3 architecture and the eight-GPU execution path. The section numbers show where each part of the layer is covered.

We use eight-way tensor parallelism for attention and MoE, with no expert parallelism (attention TP8; MoE TP8/EP1). Attention heads and each expert’s intermediate dimension are split across eight GPUs. Each GPU computes a partial result, and collectives combine the parts before dependent work begins. All benchmarks in this post were run on MI355X (CDNA4, gfx950).

Component Optimizations and Performance

Each component below begins with its role in the model, then covers kernel optimization, fusion, and distributed work where they apply. Performance numbers name their baseline and hardware.

1. KDA: Keeping Recurrent State on Chip

Standard attention stores information for every earlier token and reads that history for each new token, so its work grows with conversation length. KDA carries a fixed-size recurrent state instead. Each token updates the state and reads from it. The KDA recurrence therefore has constant work per decode step as context grows; the model’s periodic MLA layers still read the full history (Figure 3).

Figure 3: Standard attention reads a growing history; KDA updates a fixed-size recurrent state.

1.1 State layout and prefill data movement

KDA does little arithmetic per byte, so speed comes down to data movement. The recurrence computes one output value by reducing its state row across the key channels. We therefore store the state V-major as [value, key], making each reduction row contiguous. This benefits both the prefill scan and decode. The decode state uses the order its kernel reads, halving the single-user state-update time from 10.5 to 5.2us.

For prefill, the prompt is divided into 64-token chunks. Most chunk-local matrix work runs in parallel, while a scan carries the recurrent state from one chunk to the next. Our Gluon scan keeps that state in registers and produces the output in the same pass, rather than writing a checkpoint after every chunk.

Separately from the V-major state layout, we reorganize the temporary gated-key buffer used by prefill. We change it from token-major to chunk-major order, placing the scan’s reduction dimension contiguously in memory. A device-side chunk index removes per-chunk address work, and the scheduled load/MFMA pipeline overlaps memory movement with computation. These reduce the complete 8K prefill call from 659 to 595µs, a 9.7% reduction.

1.2 Fused decode and speculative state updates

A decode step originally launched separate convolution, recurrence, and normalization kernels, each writing data for the next to read. The fused Gluon kernel keeps those intermediates on chip. It also consumes packed and token-strided projection outputs directly, avoiding extra QKV and gate copies around the decode kernel. The f_b decay projection is also computed inside the recurrence kernel, removing one more launch while retaining the required BF16 rounding before gate processing. Speculative verification avoids saving a state for every rejected position; a replay step commits the accepted prefix.

1.3 KDA prefill, decode, and verification gains

  • 1.6x faster Gluon chunked prefill than the open-source Triton KDA kernel (8K tokens, 1055 -> 659us)
  • 1.11x faster KDA prefill call time at 8K after further reorganizing the gated-key buffer (8K tokens, 659- > 595us)
  • 1.63x faster decode step after fusing four kernels into one (16 users; 25.4 → 15.6 µs per layer)
  • 1.16 - 1.59x faster ordinary KDA decode including f_b projection at B1–64.
  • 1.66x less time checking speculative guesses across all 69 KDA layers (16 users, 4-token window, 2.75 → 1.66ms)

2. MLA: reusing compressed history

K3’s remaining 24 attention layers use MLA. Instead of keeping separate keys and values for 96 attention heads, MLA stores a shared 512-wide latent and a 64-wide key component per token, far less data than an expanded per-head cache. Decode can transform the query into the compressed space and project the result afterwards. Prefill may expand cached latents for its selected attention kernel. K3 uses NoPE (no rotary position embedding), and FP8 cache storage halves the bytes relative to BF16 (Figure 4).

Figure 4: MLA stores a compressed representation of the attention history. Bars drawn to scale. A smaller cache means more users and longer contexts fit in GPU memory, and less data to read per step.

2.1 Pipelined prefill and shared KV reads

For long prompts, MLA is a big compute job. Our FP8 Gluon kernel runs two groups of threads on each compute unit, offset by one step: while one group multiplies, the other loads the next KV tile in parallel.

2.2 Fused cache writes and output processing

Before an MLA kernel runs, the runtime prepares the cache rows and the query. In the prefill path, one kernel packs and converts the K/V operands and handles the cache preparation, replacing eight small launches with one. The same fused query-preparation path is used across MLA phases for normalization and projection. On pure decode, it also prepares the absorbed query. The decode epilogue then combines attention-result reduction, latent-to-value projection, output gating and the final write, avoiding a write and reread of the latent intermediate.

2.3 MLA prefill and verification gains

Figure 5: MLA prompt processing (FP8, 12 heads per GPU). time per call in µs, lower is better · measured on MI355X

3.8–5.1x faster than the portable Triton MLA kernel on full 4K and 8K chunks, and 3.2× on an 8K query with a 32K cached prefix. The 8-wave Gluon implementation is 1.13–1.18× faster than the previous Gluon kernel in these historical measurements (Figure 5).

Figure 6: Checking speculative guesses: reading the history once vs once per position. speedup, higher is better, by number of users (C) · measured on MI355X

1.86x faster at 50K context and 2.14x at 100K context for 16 users, while meeting the numerical comparison used in the kernel tests. Very short histories are slower in this experiment (0.99–1.20x at 1K/8K for 1–16 users). These are historical kernel timings, not a new end-to-end result at the cutoff (Figure 6).

3. AttnRes: fusing representation mixing

In most models, each layer’s output is added to a running residual stream. Kimi K3 groups layers in blocks of 12 and retains completed block representations. Before each attention and feed-forward step, it scores those representations together with the current accumulated stream and feeds their weighted blend into the next operation. Deep layers can select information from earlier blocks (Figure 7).

Figure 7: Attention Residuals blend saved block representations with the current residual stream.

The catch: done naively, this means rereading up to nine large tensors twice per layer. And it sits right after the tensor-parallel all-reduce, so every microsecond it takes adds directly to the step time.

3.1 Single-pass mixing with bounded register use

Our AttnRes kernel scores, blends and normalizes in a single pass over the snapshots, accumulating in 32-bit floats. We also fixed a subtle problem. The kernel used to treat one input size as a constant baked into the compiled program, so every new prompt length triggered a fresh compile in the middle of serving, at 100+ ms each. Making it a runtime value cut one long-prompt test from 46s to 16s.

The cutoff also changes the snapshot loop. Reading one snapshot at a time prevents the compiler from keeping all candidates live together, reducing register pressure. This makes an 8K-token mix of 8 snapshots 1.10× faster (283 → 257 µs). Token counts and batch-dependent strides remain runtime inputs so new serving shapes do not force compilation.

3.2 Sharing AttnRes work across GPUs

For eligible TP8 batches, we distribute the post-attention mix across token rows. Pull reduce-scatter gives each GPU one eighth of the reduced attention output. It updates the residual, mixes all saved snapshots for its assigned tokens, and applies output RMSNorm. Push all-gather then assembles the normalized MoE input on every GPU, while the updated residual shard stays local until the MoE tail consumes it (see Section 6).

This removes seven eighths of the duplicated post-attention mixing and normalization work. The same row-count and compatibility checks apply to prefill, decode, and mixed batches. For example, a compatible 64-row EAGLE3 verification batch gives each GPU eight rows to mix.

3.3 Preparing scores and fusing the final mix

This is the clearest example of why owning every kernel matters. Snapshots only change every 12 layers, so most of the blend can be computed early, as a by-product of an earlier matrix multiply. Only the last step, folding in the newest total, has to wait. We then fold that last step into the all-reduce itself (Figure 8):

Figure 8: Fusing the all-reduce, residual update, AttnRes blend, and normalization removes intermediate memory traffic.

3.4 AttnRes kernel and serving gains

Figure 9: AttnRes kernel vs a straightforward PyTorch version. time per call in µs, log scales, lower is better · measured on MI355X

4.1–5.9x faster at decode sizes and 23x faster on an 8K-token prompt chunk (Figure 9).

  • −27% time for the fused all-reduce + AttnRes step at 16 tokens (21.7 → 15.8 µs; 8 GPUs)
  • +7–14% end-to-end tokens/s for 1–4 users from that fusion and related changes

4. Latent MoE: reducing expert and projection overhead

The feed-forward network uses 896 routed experts in 92 of K3’s 93 decoder layers; the first layer is dense. A router chooses 16 experts per token, and two shared experts runs on every token. Routed experts operate in a 3,584-wide latent space rather than the 7,168-wide hidden space, then project their combined result back to full width. MXFP4 expert weights reduce storage and weight traffic; the routed and shared branches retain their distinct projections (Figure 10).

Figure 10: Latent MoE routes each token to selected experts. The grid is a sketch (224 of the 896 experts); the 16 highlighted cells are the experts chosen for one token.

4.1 Expert-weight reuse and smaller compute tiles

With 896 experts, most experts receive only a handful of tokens in any step. GPU matrix-multiply units like tidy blocks of 64 or 128 rows, so the old kernels padded each expert's few tokens up to a full block and wasted most of the work. We attacked this from both ends:

  • Decode: for small batches, walk each token’s selected routes directly instead of building sorted groups. Activations use FP8 with MXFP4 weights. At 32–64 rows in the supported K3 configuration (attention TP8; MoE TP8/EP1), we introduce moe_sorting for warp-decode MoE that sorts routes by expert so several routes reuse each loaded weight panel.
  • Prefill: use 16- and 32-row expert tiles where sparse routing would waste a larger tile, and fuse the expert output combine when the activation format and batch size make it beneficial.

4.2 Joining routed and shared expert reductions

Each GPU ends the MoE expert computation step with partial routed and shared expert results that must be summed across all eight GPUs. Our kernels write routed and shared BF16 partials directly into one reusable Iris buffer. One collective reduces both branches across eight GPUs, keeping their results separate and avoiding a concatenation copy. Tiny batches use a push protocol whose arriving data also signals readiness (Section 6).

For eligible larger batches, reduce-scatter gives each GPU both reduced results for one eighth of the token rows. It applies routed RMSNorm and up-projection locally, then adds the shared result and residual during push all-gather. Each GPU performs only one eighth of the normalization and projection work, and gathering only the combined output cuts all-gather traffic by a third.

4.3 Combining input projections and output additions

Each MoE layer applies three projections to the same input: router scores, routed latent inputs, and shared-expert gate/up values. We store their weight matrices consecutively and compute them with one matrix multiply. Where beneficial, its epilogue also applies the shared experts’ activation, avoiding an intermediate write and read.

Output fusion follows the selected expert computation. Some small-batch kernels combine the selected experts’ contributions directly; the expert-sorted 32–64-row path writes route partials and sums them in FP32. At the MoE tail, the supported 2–32-row up-projection kernels also add the shared-expert output and running residual, saving a separate addition kernel.

4.4 Expert, projection, and communication gains

Figure 11: Expert computation before and after optimization, by number of tokens. Time per call in µs, lower is better · measured on MI355X at 7e2f1deb

Together, these changes make expert computation on one GPU 2.2–2.3x faster at 32–64 tokens and 1.5–2.0x faster on 128–2,048-token prompt chunks. At 4,096–8,192 tokens, the gain is 3–4%. The comparison switches the chapter’s expert-computation optimizations off on the same code revision, isolating their combined effect (Figure 11).

Figure 12: MoE input step: one packed kernel vs three GEMMs plus SiTU. Time per call in µs, lower is better · measured on MI355X at 7e2f1deb

2.0–3.0x faster from 1 to 8,192 tokens, including 2.26x at one token and 2.95× at 8,192 tokens. The previously reported end-to-end result took single-user generation from 64.1 to 77.6 tokens/s (+21%, reported) (Figure 12).

Projection work also benefits from the Gluon GEMM additions. Bucketed Gluon decode GEMM work reports up to 3.02× faster MLA kv_b decode projection than torch.mm at a measured TP8 shape. The large-M Gluon GEMM implementation extends the large-M GEMM to ragged prefill widths and routes the measured K3 projection set. Across 160 held-out projection cases, dispatch reduces total GEMM time by about 4%. Latent up-projection fusion reduces latent up-projection plus residual-add time by 11–16% at 2–32 rows.

5. EAGLE3: verifying several tokens per pass

Speculative decoding uses a small, fast draft model to propose several tokens, then asks the target model to verify them in one pass. In the three-step EAGLE3 configuration illustrated in Figure 13, Kimi K3 checks four positions, accepts the matching prefix, and supplies a target token at the rejection point or after all proposals are accepted. Under greedy sampling, acceptance checks the target's greedy choice. Exact token equality across different batching or numerical configurations requires a separate end-to-end check.

Figure 13: EAGLE3 checks several draft tokens in one target-model pass.

Speculation changes the shape of every component's work. For example, with 16 concurrent requests and four verify positions, the target receives 64 rows. The corresponding optimizations are described where they happen: KV reuse across query positions (Section 2.1), expert-weight reuse and top-k reduction (Section 4.1– Section 4.2), and attention row sharding (Section 6.2). The three-step configuration described here uses a chain (topk=1); tree support merged by the cutoff does not change this configuration.

5.1 Performance

We evaluate EAGLE3 performance on a synthetic random dataset with 50,000 input tokens and 500 output tokens per request, comparing execution with and without EAGLE3 (Figure 14). This is a separate evaluation from the multi-turn SWE-smith workload used in the end-to-end serving comparison. EAGLE3 reduces time per output token at every tested concurrency: from 12.9 to 4.2 ms at one concurrent user, and from 28.8 to 11.4 ms at 16 concurrent users.

Figure 14: Time per output token (TPOT) with and without EAGLE3 on eight MI355X GPUs, using a synthetic random dataset with 50K input tokens and 500 output tokens per request. Lower is better; C1–C16 denote concurrent users.

6. Iris: combining communication and computation

Tensor parallelism splits each layer across eight MI355X GPUs, so their partial outputs must be summed before dependent work begins. Our Kimi K3 implementation uses Iris to build communication protocols in native Gluon and fuse computation where relevant. The symmetric heap provides buffers that kernels can directly read or write on peer GPUs. Across low latency decoding and high throughput prefill, we implemented optimized kernels based on message size and the K3 specific computation surrounding each reduction.

Each layer has two all-reduces: one after the attention output projection and one after expert computation. On the attention side, attention residuals mixes the current residual with saved states from earlier layer blocks using learned, token-dependent weights, followed by RMSNorm. The MoE all-reduce combines two partials. The routed experts produce a 3,584-wide latent vector, which we normalize and project to the model's 7,168-wide hidden dimension. The shared expert already produces a full-width output. We place both partials consecutively in one input buffer, giving a combined BF16 workload of 21 KiB per row.

6.1 Small/Medium messages

On the attention side, eligible layers fuse one-shot push all-reduce with the residual update, AttnRes combine, and final RMSNorm. Fusing these steps removes separate epilogue launches and intermediate memory traffic, reducing the fixed overhead that matters most for small messages. On the MoE side, we implemented a barrier-free all-reduce that is 1.2-3× faster than RCCL across the message sizes for which it is in use. Arriving payloads signal readiness, allowing each tile to reduce once its inputs arrive without a separate barrier. Medium-sized messages use an optimized two-shot pull all-reduce, which also handles larger messages ineligible for the row-sharded computation discussed next (Figure 15).

6.2 Large messages

Figure 15: Replicated and row-sharded communication tails: where each GPU computes and exchanges results.

Normally, each TP rank repeats the computation after all-reduce over all M. For larger messages (M >= 40 for MoE, M >= 56 for attention), we use pull reduce-scatter, assign M/8 complete rows to each of the eight GPUs, and delay push all-gather until after the attention or MoE tail. These tails communicate BF16 activations; reductions and normalization use FP32 arithmetic internally.

6.3 Collective latency and serving gains

Figures 16 and 17 compare Iris with RCCL baselines for the MoE and attention communication tails on eight MI355X GPUs.

Figure 16: MoE communication-tail latency: Iris and RCCL baselines on eight MI355X GPUs.

Conclusion

Acknowledgements

Footnotes

System configuration

What this article says