What Is Prompt Caching? How Shared KV Cache Memory Cuts AI Inference Costs

Prompt caching reuses the attention state a model has already computed for the unchanging parts of a prompt, so a GPU stops rebuilding the same system prompt, tool schema, or document on every request. Send a production AI system two requests that share the same long system prompt, even seconds apart, and there is a good chance it recomputes that prompt from scratch both times. That waste sits quietly inside a metric most teams already track: time to first token. The guide below explains what prompt caching is, how it works underneath a request, why it breaks down once you are running more than a single GPU, and what to look for in a shared context memory tier built to fix that.

What Is Prompt Caching?

Prompt caching is a technique that reuses previously computed attention state, the key-value (KV) cache, for the parts of a prompt that do not change between requests. The parts of a prompt that repeat across calls, a system prompt, a tool schema, or a long document a user keeps asking about, only get processed once instead of every time.

Think of it like a librarian who remembers exactly where she left off in a long reference book, instead of re-reading the whole book every time you walk back in with another question. The analogy also explains the limits: she only saves you time if you’re asking about the same book, starting from the same pages in the same order, before she reshelves it.

Prompt Caching vs. Prefix Caching vs. KV Cache

These three terms get used interchangeably, but they describe different layers of the same idea.

  • KV cache: the stored attention state, the key and value tensors, that a transformer model computes for every token it processes.
  • Prefix caching: the specific technique of reusing KV cache for a shared prefix, the identical leading portion of a prompt, across multiple requests.
  • Prompt caching: the product-level feature, exposed by an inference engine or API provider, that applies prefix caching (and sometimes broader reuse) automatically so the caller does not have to manage it directly.

In short: KV cache is the data, prefix caching is the mechanism, and prompt caching is what you actually turn on.

What Prompt Caching Is Not

Getting the definition right matters, because the term gets stretched to cover things it is not.

  • Not a response cache. A response cache stores a finished answer and replays it verbatim for an identical question. Prompt caching reuses intermediate computation and still generates a new response every time, which can differ even for a repeated prompt.
  • Not a vector database. A vector database retrieves semantically similar content from a large corpus. Prompt caching has nothing to do with retrieval; it reuses computation the model already performed on content already in the current prompt.
  • Not application memory. Remembering that a user prefers metric units, or what they asked last week, is a product feature built on top of a model. Prompt caching is an inference-layer optimization underneath that. The application never reads or writes the cache, but how it assembles a prompt decides whether the cache can be hit.
  • Not fuzzy matching. It only helps when a meaningful share of tokens repeat, token for token, from one request to the next. Reorder a prompt, change a timestamp inside it, or vary a system message per user, and the cache stops matching.

How Prompt Caching Works

Every request to a large language model starts with prefill: the model reads the entire input prompt and builds the KV cache for it before generating a single output token. Prefill cost scales with prompt length, so a long, mostly repeated system prompt is expensive to redo on every call, even before the model has produced anything a user can see.

Prompt caching skips that recomputation for the parts that already match. When a new request arrives, the serving engine checks how much of its prefix matches a previously cached prefix. Whatever matches gets loaded from cache instead of recomputed; only the new, non-matching tail of the prompt goes through prefill. Most serving engines match cached prefixes in fixed-size blocks of tokens rather than token by token. A block that only partially matches is recomputed, so the benefit arrives in block-sized steps up to the first point where the prompt diverges.

Decode, the phase that produces output tokens one at a time, costs the same per token whether or not the prefix was cached. What changes is how much GPU time is left for decode once prefill stops consuming it. In a busy fleet, prefill competes with decode for the same GPUs, so removing repeated prefill also raises aggregate throughput and steadies time per output token. Time to first token is the metric that moves first, not the only one that moves.

That gap alone often determines whether a system feels instant or sluggish to the person waiting on it.

Why It Matters for Enterprise AI

Two numbers turn this from a curiosity into a margin problem: time to first token, and GPU utilization.

Time to first token (TTFT) is how long a user waits before anything appears on screen. AI products have shifted from single-turn chatbots toward long-running agents with large tool schemas and long conversation histories. Along the way, TTFT has become a first-class user-experience metric rather than a background implementation detail buried in a logs dashboard.

GPU utilization is the other side of the same coin. A GPU spending cycles recomputing a system prompt it has already processed a thousand times is a GPU not doing new work, and a utilization dashboard cannot tell the two apart.

Put together, the same architectural gap shows up on a cost report as wasted GPU-hours and on a support ticket as a slow first response. That is why prompt caching gets budget attention that a pure code-quality fix usually does not.

Designing Cache-Friendly Prompts

Whether prompt caching actually helps you is mostly determined by how a prompt is assembled, not by which model you are calling.

  • Put stable content first. System prompts, tool definitions, and shared instructions belong at the beginning of a prompt; anything that changes per request, like a timestamp or a user’s latest message, belongs at the end.
  • Avoid timestamps and request IDs in shared sections. A single changing token near the front of a prompt invalidates the cache for everything that follows it.
  • Keep tool schemas and system instructions byte-identical across calls. Even whitespace or ordering differences break a prefix match, silently, with no error to alert you.
  • Batch or group similar requests where possible. Cache hit rates improve when requests sharing a prefix arrive close together, before that prefix gets evicted from memory.

Common Mistakes That Break Cache Hits

Even teams that know the rules above run into a handful of recurring mistakes.

  • Injecting a timestamp for logging purposes directly into a shared system prompt, rather than attaching it as metadata outside the cached portion.
  • Randomizing tool or function ordering on each request, which changes byte content even when the underlying capabilities have not changed.
  • Assuming caching is automatic across every deployment path, including ones that bypass the serving engine’s normal request flow, such as certain batch or evaluation pipelines.

Where Local Caching Breaks Down at Fleet Scale

A single-GPU demo makes prompt caching look simple. Production makes it hard, for three compounding reasons.

Requests Do Not Land on the Same GPU

Load balancers route requests across a fleet of replicas for availability and throughput. If a cache lives only in the memory of the specific GPU that handled the first request, every other replica has to recompute the same prefix from nothing. Session pinning and prefix-aware routing narrow the gap by steering a request back to the replica that already holds its prefix, and production gateways do this. Pinning trades away scheduling flexibility, creates hot spots, and does nothing once the prefix has been evicted or the replica has restarted. The caching benefit disappears at exactly the scale where it matters most.

HBM Is Small and Expensive

The high-bandwidth memory (HBM) a KV cache occupies during inference is the same memory the model’s weights and active batch need. A cache that competes with the model itself for HBM gets evicted quickly, especially under concurrent load from many simultaneous users.

Restarts and Autoscaling Erase Local State

Container restarts, rolling deployments, and autoscaling events all wipe whatever cache was sitting in a single node’s memory. A cache that lives and dies with one process is a single-node optimization, not fleet infrastructure.

The Shared Context Memory Tier

The fix for all three problems above is the same: move the KV cache out of any single GPU’s local memory and into a dedicated, shared tier that every replica in the fleet can reach.

A shared context memory tier is memory, not storage, even though it usually runs on infrastructure that resembles storage. Its job is to sit in the latency path of every inference request and hand back cached KV blocks fast enough that a GPU would rather wait for them than recompute them. It has to do that at the concurrency of an entire fleet, not one node. That is a different design point than durability-first object storage, even when both run on similar underlying hardware.

How This Fits the Broader Ecosystem

Prompt caching is not a single-vendor idea. Three layers of the inference stack cooperate to solve the same shared-context problem, and knowing where each one sits matters before choosing any of them.

Approach What it is Where it lives
Serving-engine prefix cache Reuses KV blocks for matching prefixes inside one engine process; lives in GPU HBM and is lost on eviction or restart vLLM automatic prefix caching, SGLang RadixAttention
KV offload and cache layer Moves KV state beyond HBM to host DRAM, local disk, or a remote backend, and defines how the engine stores and retrieves it LMCache, SGLang HiCache, Mooncake
Shared context memory tier The fleet-wide tier the offload layer writes to and reads from, reachable by every replica and persistent across restarts MinIO MemKV

The layers are complementary. The engine decides what to cache, the offload layer defines how KV state leaves the GPU and comes back, and the shared tier decides where it lives in between and who can reach it.

Evaluating a Shared Context Memory Tier

Before adopting any shared context memory layer, ask questions that separate a real fleet-scale tier from a local cache with a new name on it.

  • Does it survive a restart? If a pod restart or autoscaling event clears the cache, it is not solving the fleet-scale problem.
  • Can every replica reach it, not just the one that created it? A cache that is not shared across the fleet just relocates the local-cache problem instead of fixing it.
  • How fast does it restore a full context under concurrent load? A single offloaded block can weigh hundreds of megabytes, so restore time is a bandwidth question as much as a latency one. A shared tier only helps if restoring a context is reliably faster than recomputing it while many replicas are asking at once. Hold it to P99 time to first token on returning turns, measured with the engine's own cache left on.
  • How does it behave when it is full? Eviction policy matters as much as capacity; ask what gets dropped first and whether that matches your actual traffic pattern.
  • Does it require rewriting your serving stack? A tier that plugs into the connector interfaces your engine already supports is a materially lower-risk adoption than one that does not.

MinIO MemKV

MinIO built MemKV as a purpose-built shared context memory tier for the fleet-scale problem described above. It is designed to hold KV cache outside any single GPU’s HBM, stay reachable by every replica in a serving fleet, and survive restarts and autoscaling events.

MemKV is designed to plug into the connector interfaces used by common serving engines, including LMCache-compatible integrations, rather than requiring a rewritten serving stack.

MinIO has published two MemKV results worth reading together. The launch benchmark, Llama 3.1 70B at 64K context and production concurrency, measured time to first token at 53 seconds with full prefill recompute and 703 milliseconds with MemKV restoring the context. At 128K context the configuration without MemKV ran out of memory, while MemKV kept serving. A later benchmark built around real agentic session behavior, with the engine's own prefix cache left on and human-paced pauses between turns, measured 2.4× the tokens per second per GPU and 2.59× the completed turns on the same NVIDIA H200 GPUs (Llama-3.1-8B, 256 sessions), and a 4.5× improvement in tail latency on returning turns (Qwen3-32B, 128 sessions). The first shape of result measures a component. The second measures a fleet. Both carry their test conditions, and both are proof points to reproduce on your own traffic rather than planning inputs.

Frequently Asked Questions

Q: Does prompt caching change the model’s output? A: No. It reuses computation, not the decision of what to generate. The same prompt should produce the same distribution of possible outputs whether or not its KV cache was reused.

Q: Is prompt caching the same thing as RAG? A: No. Retrieval-augmented generation (RAG) fetches relevant content from an external corpus to add to a prompt. Prompt caching is unrelated to retrieval; it reuses computation the model already performed on content already inside the prompt.

Q: How is a shared context memory tier different from a vector database? A: A vector database stores embeddings for similarity search. A shared context memory tier stores KV cache tensors for reuse during inference. The two solve different problems and are often used in the same system without overlapping.

Q: Does prompt caching work across different models or providers? A: Generally, no. KV cache is specific to a model’s architecture and weights, so a cache built for one model cannot be reused by a different model, even from the same provider.

Q: What happens when a request does not hit the cache? A: It falls back to a normal prefill for the tokens that did not match. The added cost is the lookup itself: hashing the prompt's blocks and checking the cache, and for a remote tier, the round trip to learn the blocks are not there. Well-designed systems keep that overhead small relative to prefill, but it is not zero, which is why lookup cost belongs in the evaluation checklist above.

Conclusion

Prompt caching turns a hidden inference cost, repeatedly recomputing context, into a fleet-scale infrastructure decision. The technique itself is simple. The hard part is making it work reliably past a single GPU, which depends on where the cache lives and whether every replica in your fleet can reach it.

Request a free trial of MinIO AIStor to explore the broader data and memory platform, or talk to the MinIO team directly about running MemKV in your inference fleet.

‍

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript