
Prompt caching reuses the attention state a model has already computed for the unchanging parts of a prompt, so a GPU stops rebuilding the same system prompt, tool schema, or document on every request. Send a production AI system two requests that share the same long system prompt, even seconds apart, and there is a good chance it recomputes that prompt from scratch both times. That waste sits quietly inside a metric most teams already track: time to first token. The guide below explains what prompt caching is, how it works underneath a request, why it breaks down once you are running more than a single GPU, and what to look for in a shared context memory tier built to fix that.
Prompt caching is a technique that reuses previously computed attention state, the key-value (KV) cache, for the parts of a prompt that do not change between requests. The parts of a prompt that repeat across calls, a system prompt, a tool schema, or a long document a user keeps asking about, only get processed once instead of every time.
Think of it like a librarian who remembers exactly where she left off in a long reference book, instead of re-reading the whole book every time you walk back in with another question. The analogy also explains the limits: she only saves you time if you’re asking about the same book, starting from the same pages in the same order, before she reshelves it.
These three terms get used interchangeably, but they describe different layers of the same idea.
In short: KV cache is the data, prefix caching is the mechanism, and prompt caching is what you actually turn on.
Getting the definition right matters, because the term gets stretched to cover things it is not.
Every request to a large language model starts with prefill: the model reads the entire input prompt and builds the KV cache for it before generating a single output token. Prefill cost scales with prompt length, so a long, mostly repeated system prompt is expensive to redo on every call, even before the model has produced anything a user can see.
Prompt caching skips that recomputation for the parts that already match. When a new request arrives, the serving engine checks how much of its prefix matches a previously cached prefix. Whatever matches gets loaded from cache instead of recomputed; only the new, non-matching tail of the prompt goes through prefill. Most serving engines match cached prefixes in fixed-size blocks of tokens rather than token by token. A block that only partially matches is recomputed, so the benefit arrives in block-sized steps up to the first point where the prompt diverges.
Decode, the phase that produces output tokens one at a time, costs the same per token whether or not the prefix was cached. What changes is how much GPU time is left for decode once prefill stops consuming it. In a busy fleet, prefill competes with decode for the same GPUs, so removing repeated prefill also raises aggregate throughput and steadies time per output token. Time to first token is the metric that moves first, not the only one that moves.
That gap alone often determines whether a system feels instant or sluggish to the person waiting on it.
Two numbers turn this from a curiosity into a margin problem: time to first token, and GPU utilization.
Time to first token (TTFT) is how long a user waits before anything appears on screen. AI products have shifted from single-turn chatbots toward long-running agents with large tool schemas and long conversation histories. Along the way, TTFT has become a first-class user-experience metric rather than a background implementation detail buried in a logs dashboard.
GPU utilization is the other side of the same coin. A GPU spending cycles recomputing a system prompt it has already processed a thousand times is a GPU not doing new work, and a utilization dashboard cannot tell the two apart.
Put together, the same architectural gap shows up on a cost report as wasted GPU-hours and on a support ticket as a slow first response. That is why prompt caching gets budget attention that a pure code-quality fix usually does not.
Whether prompt caching actually helps you is mostly determined by how a prompt is assembled, not by which model you are calling.
Even teams that know the rules above run into a handful of recurring mistakes.
A single-GPU demo makes prompt caching look simple. Production makes it hard, for three compounding reasons.
Load balancers route requests across a fleet of replicas for availability and throughput. If a cache lives only in the memory of the specific GPU that handled the first request, every other replica has to recompute the same prefix from nothing. Session pinning and prefix-aware routing narrow the gap by steering a request back to the replica that already holds its prefix, and production gateways do this. Pinning trades away scheduling flexibility, creates hot spots, and does nothing once the prefix has been evicted or the replica has restarted. The caching benefit disappears at exactly the scale where it matters most.
The high-bandwidth memory (HBM) a KV cache occupies during inference is the same memory the model’s weights and active batch need. A cache that competes with the model itself for HBM gets evicted quickly, especially under concurrent load from many simultaneous users.
Container restarts, rolling deployments, and autoscaling events all wipe whatever cache was sitting in a single node’s memory. A cache that lives and dies with one process is a single-node optimization, not fleet infrastructure.
The fix for all three problems above is the same: move the KV cache out of any single GPU’s local memory and into a dedicated, shared tier that every replica in the fleet can reach.
A shared context memory tier is memory, not storage, even though it usually runs on infrastructure that resembles storage. Its job is to sit in the latency path of every inference request and hand back cached KV blocks fast enough that a GPU would rather wait for them than recompute them. It has to do that at the concurrency of an entire fleet, not one node. That is a different design point than durability-first object storage, even when both run on similar underlying hardware.
Prompt caching is not a single-vendor idea. Three layers of the inference stack cooperate to solve the same shared-context problem, and knowing where each one sits matters before choosing any of them.
The layers are complementary. The engine decides what to cache, the offload layer defines how KV state leaves the GPU and comes back, and the shared tier decides where it lives in between and who can reach it.
Before adopting any shared context memory layer, ask questions that separate a real fleet-scale tier from a local cache with a new name on it.
MinIO built MemKV as a purpose-built shared context memory tier for the fleet-scale problem described above. It is designed to hold KV cache outside any single GPU’s HBM, stay reachable by every replica in a serving fleet, and survive restarts and autoscaling events.
MemKV is designed to plug into the connector interfaces used by common serving engines, including LMCache-compatible integrations, rather than requiring a rewritten serving stack.
MinIO has published two MemKV results worth reading together. The launch benchmark, Llama 3.1 70B at 64K context and production concurrency, measured time to first token at 53 seconds with full prefill recompute and 703 milliseconds with MemKV restoring the context. At 128K context the configuration without MemKV ran out of memory, while MemKV kept serving. A later benchmark built around real agentic session behavior, with the engine's own prefix cache left on and human-paced pauses between turns, measured 2.4× the tokens per second per GPU and 2.59× the completed turns on the same NVIDIA H200 GPUs (Llama-3.1-8B, 256 sessions), and a 4.5× improvement in tail latency on returning turns (Qwen3-32B, 128 sessions). The first shape of result measures a component. The second measures a fleet. Both carry their test conditions, and both are proof points to reproduce on your own traffic rather than planning inputs.
Q: Does prompt caching change the model’s output? A: No. It reuses computation, not the decision of what to generate. The same prompt should produce the same distribution of possible outputs whether or not its KV cache was reused.
Q: Is prompt caching the same thing as RAG? A: No. Retrieval-augmented generation (RAG) fetches relevant content from an external corpus to add to a prompt. Prompt caching is unrelated to retrieval; it reuses computation the model already performed on content already inside the prompt.
Q: How is a shared context memory tier different from a vector database? A: A vector database stores embeddings for similarity search. A shared context memory tier stores KV cache tensors for reuse during inference. The two solve different problems and are often used in the same system without overlapping.
Q: Does prompt caching work across different models or providers? A: Generally, no. KV cache is specific to a model’s architecture and weights, so a cache built for one model cannot be reused by a different model, even from the same provider.
Q: What happens when a request does not hit the cache? A: It falls back to a normal prefill for the tokens that did not match. The added cost is the lookup itself: hashing the prompt's blocks and checking the cache, and for a remote tier, the round trip to learn the blocks are not there. Well-designed systems keep that overhead small relative to prefill, but it is not zero, which is why lookup cost belongs in the evaluation checklist above.
Prompt caching turns a hidden inference cost, repeatedly recomputing context, into a fleet-scale infrastructure decision. The technique itself is simple. The hard part is making it work reliably past a single GPU, which depends on where the cache lives and whether every replica in your fleet can reach it.
Request a free trial of MinIO AIStor to explore the broader data and memory platform, or talk to the MinIO team directly about running MemKV in your inference fleet.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Block quote
Ordered list
Unordered list
Bold text
Emphasis
Superscript
Subscript