Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/How to integrate lmcache with flashblade for faster llm inference
Dev48

© 2026 · All rights reserved.

How to Integrate LMCache with FlashBlade for Faster LLM Inference

Источник: Everpure Blog

How to Integrate LMCache with FlashBlade for Faster LLM Inference

Source: Everpure Blog

How to Integrate LMCache with FlashBlade for Faster LLM Inference by Everpure Blog Explore how LMCache and Everpure FlashBlade can enable persistent KV caching to help reduce time to first token and increase LLM inference throughput. The post How to Integrate LMCache with FlashBlade for Faster…

September 29, 2026•Updated: September 29, 2026

Introduction

A growing number of our customers want ownership of their AI operations and data, which in practice means running open-weight models on infrastructure they own with proprietary data that never leaves their walls.

When running your own models on your own GPUs, maximizing utilization is paramount. When it comes to running inference, loading an already saved KV cache is an easy way to accomplish this. To learn more about KV caching, see this blog. To summarize, for each new LLM input/query, we save the computed KV vectors from the prefill. Then, when a similar query hits the same prefix, we load the relevant KV cache, freeing the GPU to compute new, unseen queries instead of recomputing a known entity.

This blog covers several ways Everpure storage solutions can be leveraged by LMCache.

What is LMCache?

LMCache is an open source KV cache solution that operates as a KV cache management layer designed for LLM inference. It converts ephemeral KV cache data into durable, reusable AI-native knowledge that can be stored persistently and shared across multiple serving engines. By leveraging LMCache, systems achieve lower time to first token (TTFT) and enhanced throughput, delivering significant gains for long-context workloads like multi-turn conversations, agentic tasks, and RAG.

Built with a vendor-neutral design, LMCache seamlessly integrates across a wide ecosystem of open source serving engines, inference frameworks, storage backends, hardware platforms, and infrastructure providers.

LMCache can run in two modes: in-process and multi-process.

Figure 1: LMCache modes. Source.

Everpure FlashBlade storage options for LMCache

Everpure provides several storage platforms that can be leveraged by LMCache depending on your environment and requirements.

FlashBlade//S™ is a flexible, scale-out platform for enterprise unstructured data, supporting workloads such as AI, analytics, backup, and rapid restore with independent scaling of performance and capacity.

FlashBlade//EXA™ extends the family for the most demanding GPU cloud and AI application environments, delivering extreme throughput and performance at massive scale.

LMCache and FlashBlade integration examples

The following examples show how to integrate LMCache with Everpure™ FlashBlade®.

It’s important to note that results and benefits of adding LMCache are dependent on the environment itself. Environments with different GPUs, networking, and software versions are not directly comparable.

This environment uses:

  • CPU: 1x Intel Xeon Gold 5515+, 32 cores
  • Linux kernel: 6.8.0-124
  • GPU: NVIDIA L40S, 46,068 MiB, SM 8.9
  • NICs: 1x 100 GbE for S3, 1x 100 GbE for GDS
  • NFSv3 with nconnect=16 and proto=rdma
  • lmcache==0.5.3
  • vllm==0.24.0

In-process mode with FlashBlade GPU Direct Storage (GDS)

LMCache’s GDS backend uses NVIDIA cuFile (or AMD hipFile) for zero-copy I/O between GPU memory and a mounted filesystem, avoiding the CPU bounce buffer. This is the lowest-latency path on dedicated inference nodes with Everpure FlashBlade.

To run LMCache in-process with GDS, we use the following lmcache.yaml definitions file:

We start vLLM by specifying the LMCACHE_CONFIG_FILE. Here is the full command used:

In-process mode with FlashBlade S3

LMCache ships an S3 connector, and FlashBlade serves S3 natively. For multi-tenant or multi-site fleets, object semantics are usually the easier operational answer: identity and policy per bucket rather than per mount, and no client-side mount configuration to manage across a large fleet.

To run LMCache in-process with S3, we use the following lmcache.yaml definitions file:

We start vLLM by specifying the LMCACHE_CONFIG_FILE. Here is the full command used:

vllm serve openai/gpt-oss-20b \ --port 8000 \ --tensor-parallel-size 1 \ --gpu-memory-utilization 0.92 \ --max-model-len 131072 \ --disable-hybrid-kv-cache-manager \ --kv-transfer-config '{"kv_connector":"LMCacheConnectorV1","kv_role":"kv_both"}' \ --no-enable-prefix-caching

Multi-process mode with L1 CPU and L2 FlashBlade NFS

In contrast to in-process mode, when running LMCache in multi-process mode, the lmcache process and vLLM process are separate. This allows multiple vLLM processes to leverage a shared LMCache service.

It must be noted that in these dual configurations, L1 is memory based, with L2 only used when LRU has moved the cache hit out. Due to this, we will not showcase the configuration in our results, since, in essence, L2 results are comparable to GDS in-process results.

First, we start LMCache using the lmcache server command and provide the relevant L1 and L2 settings for our integration:

Then, we start our vLLM service, with the relevant LMCache service defined:

Results

Here, we see the difference in time to first token (TTFT) for each of the different LMCache setups compared against a cold, uncached prefill (no LMCache). TTFT is a measure of the delay before the LLM starts responding and is responsible for how fast or how sluggish the model feels. Faster TTFT results in a better user experience.

Without caching, the model takes 33.5 seconds to respond with a context length of 128K. The longer the context length, the longer the response time. With LMCache on FlashBlade S3 or GDS, TTFT at a 128K context length is 3.4 and 1.9 seconds, respectively.

Figure 2: Time to first token with and without LMCache.

Below, we can see the improvement in TTFT for the different LMCache backends. FlashBlade is used with S3 and GDS (NFS over RDMA). FlashBlade supports both on the same system, so you can choose which protocol best fits your needs. Note that the improvement in TTFT increases as the context length increases.

Figure 3: Time to first token improvement with LMCache on FlashBlade S3 and GDS.

End-to-end token throughput also increases substantially when using LMCache with FlashBlade backends. The reference peaks at 32K and then declines because prefill grows faster than the tokens it produces. However, token throughput for both FlashBlade S3 and GDS continues to increase all the way to 128K context, which is where the testing stopped due to the context length of the model used.

Figure 4: End-to-end token throughput with and without LMCache.

The improvement in end-to-end token throughput also continues to increase as context length increases. Throughput gain increased to 4.97X for FlashBlade with GDS and 4.10X for FlashBlade with S3.

Figure 5: Throughput improvement with LMCache on FlashBlade S3 and GDS.

Conclusion

Whether you need the raw, zero-copy throughput of GDS on FlashBlade or the operational simplicity of S3 for multi-tenant fleets of instances, Everpure storage gives LMCache the durable, high-performance backend it needs to turn KV cache from a disposable byproduct into a reusable asset. The result is lower TTFT, higher throughput, and GPUs spent computing new work instead of recomputing what’s already been seen—all on infrastructure you own and control.

← All articles

More in Hardware & Electronics

All →
Alaska Airlines CEO 'not overly concerned' about new Boeing Max 10 delayПресса
Boeing

Alaska Airlines CEO 'not overly concerned' about new Boeing Max 10 delay

Oura shelves its $2.2B IPO citing ‘uncertainty’ in the marketПресса
Oura

Oura shelves its $2.2B IPO citing ‘uncertainty’ in the market

Meta is expanding its AI agent Muse to small businesses
Пресса
Meta

Meta is expanding its AI agent Muse to small businesses

EagleNXT Participates in Michigan NADWC DDIL Testing Environment Supporting Army Permissive Range Initiative
Ageagle

EagleNXT Participates in Michigan NADWC DDIL Testing Environment Supporting Army Permissive Range Initiative

Meta launches Muse for Small Business as Zuckerberg pushes beyond consumer AI marketПресса
Meta

Meta launches Muse for Small Business as Zuckerberg pushes beyond consumer AI market

Smart PC Building in 2026: What Users Need to Know Before They Buy
ASUS

Smart PC Building in 2026: What Users Need to Know Before They Buy

More from Pure Storage

Making Sustainability Count: Turning Environmental Impact into Business Insight
Pure Storage

Making Sustainability Count: Turning Environmental Impact into Business Insight

Everpure Brings Carrier-Grade Storage to Red Hat’s Telco Cloud Reference Architecture
Pure Storage

Everpure Brings Carrier-Grade Storage to Red Hat’s Telco Cloud Reference Architecture

Why the Everpure Growth Equation Has Changed
Pure Storage

Why the Everpure Growth Equation Has Changed

The Market Has Spoken: QLC’s Hyperscale Moment Has Arrived
Pure Storage

The Market Has Spoken: QLC’s Hyperscale Moment Has Arrived