Skip to content
All articles
  • Applied AI
  • LLM inference
  • KV cache

KV cache and why it matters for AI inference

How KV cache works, why long chats consume GPU memory, and how to estimate the number of conversations a GPU can hold.

Eduard Smirnov11 min read

A model can load on your GPU, answer a short question, and still run out of memory during a long conversation. The model's weights have not grown. One common reason is that the KV cache has.

KV means key and value. The cache stores intermediate results from the model's attention layers, so it can use the context again without calculating those results from scratch. This is a basic part of making transformer language models practical to serve.

That extra memory is why a model fitting on a GPU says little about how many long conversations it can serve.

What the model keeps in the cache

A language model works with tokens, which can be words, parts of words or punctuation. In a typical decoder transformer, attention lets the current position use information from earlier positions.

Attention produces three kinds of vectors. A query is compared with keys to determine which positions receive attention. The corresponding values are then combined using those attention weights. These are learned numerical representations, not literal questions, keywords or database records. The mechanism is described in the original Transformer paper.

During generation, the model needs the earlier keys and values repeatedly. A KV cache keeps them for each attention layer. At the next step, it calculates the new position's query, key and value, adds the new key and value to the cache, and uses the cached context to predict the following token. Hugging Face's explanation walks through this process.

Why can those older results be reused? In a causal decoder, a position cannot attend to future positions. Appending more tokens therefore does not change the earlier states for the same prefix and model configuration. Changing an earlier input is a different matter.

The cache holds attention states for a particular context. It does not train the model, add facts to its weights or store a finished answer for every possible question.

From reading the prompt to writing the answer

Suppose an assistant receives an 8,000-token maintenance manual and generates a 200-token answer.

A naive implementation without a cache could run the growing sequence through the model for every output token. The original 8,000 prompt positions would be processed 200 times, or 1.6 million prompt-position visits. With caching, their keys and values are computed once and reused. Newly processed positions still need their own calculations.

The 200-fold reduction applies to those repeated prompt calculations. The whole answer does not become 200 times faster: the model still has to attend to the context and generate each output token.

Prefill processes an 8,000-token prompt, caches its keys and values, and predicts output token 1. The next step processes token 1, reuses the prompt cache, adds one position, and predicts token 2. Processing token 2 grows the cache to 8,002 positions and predicts token 3.
A generated token enters the cache when the model processes it on the next step. The count here is the number of cached positions in each layer.

There are two phases to keep separate. Prefill processes the prompt and builds its attention states, with substantial work across prompt positions happening in parallel. Decode processes subsequent positions as the answer grows. NVIDIA's inference guide explains why they have different performance characteristics.

For full attention, a new query still uses the earlier keys and values. A longer context means more cache data to read. Decode can become limited by memory bandwidth, as the multi-query attention paper explains.

This is why a long prompt can delay the first token, while a long retained context can also affect the pace of later tokens. "Tokens per second" alone hides that distinction.

How a long chat consumes GPU memory

The memory cost depends on the architecture, cache precision and number of retained tokens. For a conventional decoder with full attention, equally sized key and value vectors, and the same cache layout in every layer, a useful estimate is:

cache bytes per sequence =
    2 * layers * KV heads * head dimension * retained tokens * bytes per element

The first 2 accounts for both keys and values. An element is one stored number in either tensor; a 16-bit element takes 2 bytes. This follows the cache-size calculation in NVIDIA's guide, with KV heads times head dimension in place of hidden size. That distinction matters for models whose query heads share keys and values.

Use Llama 3.1 8B as a concrete example. Meta's model paper lists 32 layers, 8 KV heads per layer, a model dimension of 4,096 and 32 query heads. Dividing 4,096 by 32 gives a head dimension of 128. With a 16-bit cache:

2 * 32 * 8 * 128 * 2 = 131,072 bytes per token
                          = 128 KiB per token

Counting prompt and processed output tokens together gives:

Retained tokens in one sequence Approximate KV storage
4,096 512 MiB
16,384 2 GiB
32,768 4 GiB
65,536 8 GiB

The units are binary: one GiB is 1,073,741,824 bytes. Weights and runtime memory are additional.

At 32,768 retained tokens each, 16 independent sequences would require about 64 GiB of KV storage under these assumptions, before those other costs. Sharing an identical cached prefix can reduce duplication, but unrelated conversations do not get that saving.

Here is the calculation in plain Python:

layers = 32
kv_heads = 8
head_dim = 128
bytes_per_element = 2
retained_tokens = 32_768
sequences = 16

per_token = 2 * layers * kv_heads * head_dim * bytes_per_element
per_sequence = per_token * retained_tokens
total = per_sequence * sequences

print(f"Per token: {per_token / 1024:.0f} KiB")
print(f"Per sequence: {per_sequence / 1024**3:.1f} GiB")
print(f"Total KV data: {total / 1024**3:.1f} GiB")

Leave room for the answer to grow, too. A request that fits at the end of prefill may need additional cache storage as generation continues.

How many long conversations fit on one GPU?

Use the memory total reported by nvidia-smi. NVIDIA's DGX deployment guide shows an H100 80GB reporting 81,559 MiB, about 79.6 GiB, with MIG disabled. That is the full-card total used here. Roughly 8 billion parameters at 2 bytes per bf16 weight use 14.9 GiB.

Allow another 8 GiB for runtime allocations and unused headroom combined, leaving:

79.6 GiB reported total - 14.9 GiB weights - 8 GiB allowance = 56.7 GiB for KV

At 4 GiB per sequence, that leaves room for 14 independent conversations at 32,768 retained tokens each, with a 16-bit cache and no prefix sharing. The token count includes the prompt and processed answer. Memory capacity still needs a latency test: fitting all the conversations does not establish how quickly the server can answer them.

An H100 80GB with a reported total of 81,559 MiB has about 79.6 GiB. Subtracting 14.9 GiB for bf16 weights and a combined 8 GiB runtime and headroom allowance leaves 56.7 GiB for KV cache. At 4 GiB per independent 32,768-token conversation, fourteen full conversations fit in this example.
Start with the reported memory total and count runtime use and headroom once.

The serving engine applies its own budget. vLLM's cache configuration documentation lists gpu_memory_utilization=0.92 as the default; check your installed version. This fraction covers the whole model executor. Its memory profiling code subtracts weights and runtime allocations before assigning the remainder to KV cache.

Its startup report gives the KV capacity in tokens and maximum concurrency at the configured maximum sequence length. Use that cache pool directly: the engine has already applied its budget. Subtracting the example's 8 GiB again would count overhead twice. Scheduler limits such as max_num_seqs can further restrict the active batch.

What memory bandwidth says about speed

For fourteen independent 32K conversations, assume a decode batch produces one token per chat, reading the weights once and each full KV cache once from GPU memory:

Weights: about 16 GB
KV data: 14 * 4 GiB = 56 GiB, about 60 GB
Total reads: about 76 GB per batch step
Memory time = bytes read / bytes per second

The cache now accounts for most of the traffic. NVIDIA's H100 datasheet lists 3.35 TB/s for the SXM card and 2 TB/s for the 80GB PCIe card. Using decimal GB and TB for the bandwidth calculation gives:

H100 variant Peak memory bandwidth Ideal batch step Rate per conversation
SXM 80GB 3.35 TB/s About 23 ms About 44 tokens/s
PCIe 80GB 2 TB/s About 38 ms About 26 tokens/s

These are optimistic bandwidth-only estimates for the assumed reads. Compute and kernel overhead can increase step time, and sustained bandwidth is usually below the advertised peak. Shared prefixes, lower cache precision or speculative decoding change the traffic or tokens produced per step.

Why the number of KV heads matters

Attention heads let a layer combine information in several ways. In ordinary multi-head attention, each query head has its own keys and values. Grouped-query attention, or GQA, lets several query heads share a KV head. Multi-query attention, or MQA, shares one KV head across the query heads. The GQA paper describes this design tradeoff.

In our example, changing the KV head count from 8 to 32 would multiply the cache data by four. A 32,768-token sequence would then need 16 GiB instead of 4 GiB, with all other settings held constant.

Changing that count requires a compatible checkpoint. The GQA paper shows how existing multi-head checkpoints can be converted and trained further. You cannot get the same result by changing a head-count setting at inference time. Parameter count by itself does not tell you the cache footprint.

Where the simple formula stops applying

Multi-head latent attention, or MLA, changes what is stored. DeepSeek-V2 caches a compressed representation of keys and values plus separate positional key data. Its cache still grows with retained tokens, but the per-token layout differs from the full K and V tensors counted above. The DeepSeek-V2 paper explains the compression.

Other architectures avoid keeping a separate KV entry for every past token. The Mamba paper describes a state-space model that updates a fixed-size recurrent state during generation. The Transformers are RNNs paper shows how causal linear attention can also accumulate context in a state whose storage does not grow with sequence length. That fixed-size state is a lossy summary of the past: details can be lost, making exact recall from far back in a long context harder.

Hybrid models can mix recurrent layers with attention layers. Count the storage for each kind separately: a fixed recurrent state does not stop the model's full-attention layers from accumulating KV data.

Sliding-window attention provides another limit: a layer keeps a bounded recent window instead of the entire history. Transformers documents cache behaviour for sliding-window and hybrid attention layouts. Check the model's actual layers before multiplying one full-attention estimate across all of them.

Prefix caching across requests

Prefix caching lets requests with the same initial token sequence reuse cached states while those states remain available.

Take a handbook assistant serving twenty employees. Each asks a different question, but every request starts with the same instructions and the same 12,000-token handbook. A serving engine that supports prefix reuse can retain the handbook's states after the first request and use them for later questions. vLLM's automatic prefix caching documentation describes this behaviour.

With one cold request followed by 19 complete prefix hits, processing that shared prefix falls from 240,000 token positions to 12,000. Each question's suffix still needs processing, and every answer still needs generation. Cache eviction, routing to a different worker or incomplete matches can reduce the saving.

Prompt layout makes a difference:

Useful shared prefix
  Stable instructions
  Stable handbook
  Employee's question

Prefix changes early
  Unique request ID and timestamp
  Stable instructions
  Stable handbook
  Employee's question

In the second layout, the requests diverge before the handbook. Identical document text later in the prompt does not make its attention states interchangeable, because those states depend on the preceding context. Put changing metadata later when the task allows it, and keep shared material consistently ordered.

Hosted APIs may expose this as prompt caching, with their own eligibility, retention and pricing. Ordinary generation caching does not automatically give you a cache hit or a discount on the next request.

Prefix reuse saves repeated prefill work; every answer still needs generation. That is the saving described in vLLM's documentation. Sharing cached states can also free memory for larger decode batches. Short questions about a long shared document benefit most directly from skipping its prefill.

Ways to manage cache memory

Paging addresses allocation. Without careful management, a server can waste memory on oversized reservations and fragmented storage. PagedAttention stores KV data in blocks and maps each sequence to its blocks, allowing more flexible allocation and sharing. The PagedAttention paper explains the approach behind vLLM.

It does not make every retained token smaller. Efficient placement and compression solve different problems.

For local inference, the Transformers cache guide describes several choices:

Approach Main tradeoff
Dynamic cache Grows with the sequence; changing shapes can complicate compilation.
Static cache Reserves a maximum size to support compilation; unused capacity still costs memory.
Cache offloading Moves cache data between GPU and CPU memory; transfers can reduce throughput.
Cache quantization Stores KV data at lower precision; conversion costs can hurt latency, and quality needs evaluation.

Quantizing weights and quantizing the cache are separate operations. A model loaded with 4-bit weights does not necessarily have a 4-bit KV cache. Check both configurations when estimating memory.

Token eviction reduces how many positions remain in the cache. The H2O paper describes retaining recent tokens and "heavy hitters" that receive substantial attention, while discarding others. It can bound cache growth without lowering each element's precision. Evicting tokens from an active sequence changes the context available to later attention, so test whether the model still answers your questions correctly.

That differs from releasing an unused shared prefix cache. A future prefix-cache miss can be handled by processing the full prompt again. An active sequence with discarded positions no longer has all of those keys and values available.

Reducing context is another option, but it changes what the model can use. For a maintenance assistant, removing a duplicate page is different from removing the safety exception needed to answer the question. Evaluate the resulting answers, rather than celebrating a smaller allocation on its own.

What to measure in a real application

Start with the prompt lengths, answer lengths and concurrency your application actually needs. A one-user test with a tiny prompt will miss the part of the workload where cache management matters most.

Measure time to first token separately from the time between later tokens. Record peak memory and throughput under concurrent requests. For prefix reuse, separate cold requests from warm hits and check how many input tokens were actually reused. Include misses and evictions in the test, rather than warming one prompt forever.

Test generation caching and prefix reuse separately, with each enabled and disabled. When changing cache precision or reducing context, check answer quality alongside speed and memory.

The 8B model's weights occupy about 15 GiB. Its live conversations can add 56 GiB of cache that attention has to read as answers grow. That extra memory affects both how many chats fit and how quickly each one receives a response.