Introduction
A growing number of our customers want ownership of their AI operations and data, which in practice means running open-weight models on infrastructure they own with proprietary data that never leaves their walls.
When running your own models on your own GPUs, maximizing utilization is paramount. When it comes to running inference, loading an already saved KV cache is an easy way to accomplish this. To learn more about KV caching, see this blog. To summarize, for each new LLM input/query, we save the computed KV vectors from the prefill. Then, when a similar query hits the same prefix, we load the relevant KV cache, freeing the GPU to compute new, unseen queries instead of recomputing a known entity.
This blog covers several ways Everpure storage solutions can be leveraged by LMCache.
What is LMCache?
LMCache is an open source KV cache solution that operates as a KV cache management layer designed for LLM inference. It converts ephemeral KV cache data into durable, reusable AI-native knowledge that can be stored persistently and shared across multiple serving engines. By leveraging LMCache, systems achieve lower time to first token (TTFT) and enhanced throughput, delivering significant gains for long-context workloads like multi-turn conversations, agentic tasks, and RAG.
Built with a vendor-neutral design, LMCache seamlessly integrates across a wide ecosystem of open source serving engines, inference frameworks, storage backends, hardware platforms, and infrastructure providers.
LMCache can run in two modes: in-process and multi-process.
Figure 1: LMCache modes. Source.
Everpure FlashBlade storage options for LMCache
Everpure provides several storage platforms that can be leveraged by LMCache depending on your environment and requirements.
FlashBlade//S™ is a flexible, scale-out platform for enterprise unstructured data, supporting workloads such as AI, analytics, backup, and rapid restore with independent scaling of performance and capacity.
FlashBlade//EXA™ extends the family for the most demanding GPU cloud and AI application environments, delivering extreme throughput and performance at massive scale.
LMCache and FlashBlade integration examples
The following examples show how to integrate LMCache with Everpure™ FlashBlade®.
It’s important to note that results and benefits of adding LMCache are dependent on the environment itself. Environments with different GPUs, networking, and software versions are not directly comparable.
This environment uses:
- CPU: 1x Intel Xeon Gold 5515+, 32 cores
- Linux kernel: 6.8.0-124
- GPU: NVIDIA L40S, 46,068 MiB, SM 8.9
- NICs: 1x 100 GbE for S3, 1x 100 GbE for GDS
- NFSv3 with nconnect=16 and proto=rdma
- lmcache==0.5.3
- vllm==0.24.0
In-process mode with FlashBlade GPU Direct Storage (GDS)
LMCache’s GDS backend uses NVIDIA cuFile (or AMD hipFile) for zero-copy I/O between GPU memory and a mounted filesystem, avoiding the CPU bounce buffer. This is the lowest-latency path on dedicated inference nodes with Everpure FlashBlade.
To run LMCache in-process with GDS, we use the following lmcache.yaml definitions file:
We start vLLM by specifying the LMCACHE_CONFIG_FILE. Here is the full command used:
In-process mode with FlashBlade S3
LMCache ships an S3 connector, and FlashBlade serves S3 natively. For multi-tenant or multi-site fleets, object semantics are usually the easier operational answer: identity and policy per bucket rather than per mount, and no client-side mount configuration to manage across a large fleet.
To run LMCache in-process with S3, we use the following lmcache.yaml definitions file:
We start vLLM by specifying the LMCACHE_CONFIG_FILE. Here is the full command used:
vllm serve openai/gpt-oss-20b \ --port 8000 \ --tensor-parallel-size 1 \ --gpu-memory-utilization 0.92 \ --max-model-len 131072 \ --disable-hybrid-kv-cache-manager \ --kv-transfer-config '{"kv_connector":"LMCacheConnectorV1","kv_role":"kv_both"}' \ --no-enable-prefix-caching
Multi-process mode with L1 CPU and L2 FlashBlade NFS
In contrast to in-process mode, when running LMCache in multi-process mode, the lmcache process and vLLM process are separate. This allows multiple vLLM processes to leverage a shared LMCache service.
It must be noted that in these dual configurations, L1 is memory based, with L2 only used when LRU has moved the cache hit out. Due to this, we will not showcase the configuration in our results, since, in essence, L2 results are comparable to GDS in-process results.
First, we start LMCache using the lmcache server command and provide the relevant L1 and L2 settings for our integration:
Then, we start our vLLM service, with the relevant LMCache service defined:
Results
Here, we see the difference in time to first token (TTFT) for each of the different LMCache setups compared against a cold, uncached prefill (no LMCache). TTFT is a measure of the delay before the LLM starts responding and is responsible for how fast or how sluggish the model feels. Faster TTFT results in a better user experience.
Without caching, the model takes 33.5 seconds to respond with a context length of 128K. The longer the context length, the longer the response time. With LMCache on FlashBlade S3 or GDS, TTFT at a 128K context length is 3.4 and 1.9 seconds, respectively.
Figure 2: Time to first token with and without LMCache.
Below, we can see the improvement in TTFT for the different LMCache backends. FlashBlade is used with S3 and GDS (NFS over RDMA). FlashBlade supports both on the same system, so you can choose which protocol best fits your needs. Note that the improvement in TTFT increases as the context length increases.
Figure 3: Time to first token improvement with LMCache on FlashBlade S3 and GDS.
End-to-end token throughput also increases substantially when using LMCache with FlashBlade backends. The reference peaks at 32K and then declines because prefill grows faster than the tokens it produces. However, token throughput for both FlashBlade S3 and GDS continues to increase all the way to 128K context, which is where the testing stopped due to the context length of the model used.
Figure 4: End-to-end token throughput with and without LMCache.
The improvement in end-to-end token throughput also continues to increase as context length increases. Throughput gain increased to 4.97X for FlashBlade with GDS and 4.10X for FlashBlade with S3.
Figure 5: Throughput improvement with LMCache on FlashBlade S3 and GDS.
Conclusion
Whether you need the raw, zero-copy throughput of GDS on FlashBlade or the operational simplicity of S3 for multi-tenant fleets of instances, Everpure storage gives LMCache the durable, high-performance backend it needs to turn KV cache from a disposable byproduct into a reusable asset. The result is lower TTFT, higher throughput, and GPUs spent computing new work instead of recomputing what’s already been seen—all on infrastructure you own and control.








