- PagedAttention applies virtual memory paging to KV cache, eliminating 60-80% memory waste from pre-allocation.
- The 24x throughput claim is real but workload-dependent—expect 2-4x improvement on typical chat workloads.
- Copy-on-write sharing enables efficient beam search by sharing KV blocks until sequences diverge.
- The technique adds minimal overhead (~4%) because block sizes match GPU memory transaction granularity.
- For production LLM serving, vLLM with PagedAttention is the clear choice over raw HuggingFace Transformers.
Why Memory Fragmentation Kills LLM Serving Throughput
Here’s a number that should make you uncomfortable: existing LLM serving systems waste 60-80% of their GPU memory on KV cache fragmentation. That’s not a typo. When you’re paying $3/hour for an A100, you’re effectively burning $1.80-$2.40 of that on empty memory blocks that can’t be used.
The PagedAttention paper from Kwon et al. (SOSP 2023) tackles this directly. You can read the full paper here. The headline claim—24x throughput improvement over HuggingFace Transformers—sounds like marketing hype until you understand what’s actually happening under the hood.

The KV Cache Problem Nobody Talks About
When a Transformer generates tokens, it needs to store the Key and Value vectors from all previous positions. For a 13B parameter model like LLaMA-13B, each token requires about 800KB of KV cache storage. A single sequence with 2048 tokens needs 1.6GB just for KV cache—and that’s per sequence.
The naive approach (what HuggingFace Transformers does) pre-allocates a contiguous memory block for the maximum possible sequence length. If your max length is 2048 tokens but your actual prompt is 50 tokens, you’ve just reserved memory for 1998 tokens that don’t exist yet. And might never exist if generation stops early.
This is where it gets worse. Because memory is allocated contiguously, you can’t reclaim the unused portions. You can’t pack multiple sequences efficiently. The memory sits there, reserved but empty, while other requests queue up waiting for GPU space.
PagedAttention: Borrowing From OS Virtual Memory
The core insight is almost embarrassingly simple once you see it: treat KV cache like virtual memory pages.
Instead of allocating one big contiguous block per sequence, PagedAttention divides the KV cache into fixed-size blocks (typically 16 tokens per block). Each sequence gets a block table that maps logical token positions to physical memory blocks. Just like how your operating system maps virtual addresses to physical RAM pages.
The attention computation becomes:
But now and are gathered from non-contiguous physical blocks using the block table. The kernel fetches blocks from wherever they happen to be in GPU memory, computes partial attention scores, and combines them.
This sounds like it should add overhead. Random memory access patterns are typically slower than sequential reads. But the paper shows the overhead is minimal—around 4% slowdown in the attention kernel itself—because the block size is chosen to match GPU cache lines and memory transaction sizes.
Memory Efficiency: The Real Win
With paging, memory waste drops from 60-80% to near zero. Here’s why:
- No pre-allocation: Blocks are allocated on-demand as tokens are generated
- Immediate reclamation: When a sequence finishes, its blocks return to the free pool instantly
- Copy-on-write sharing: Parallel sampling (beam search, parallel decoding) can share KV cache blocks until they diverge
The copy-on-write mechanism is particularly clever for beam search. When you have 4 beams, they all start with the same prompt. Traditional systems allocate 4 complete copies of the prompt’s KV cache. vLLM allocates one copy and has all beams point to the same physical blocks. Only when beams diverge do they get separate blocks.
For a beam search with 4 beams on a 1024-token prompt, this saves roughly 75% of the prompt’s KV cache memory. The memory savings scale linearly with beam width.
The 24x Throughput Claim: Breaking It Down
The paper’s headline number—24x improvement over HuggingFace Transformers—comes from a specific benchmark: OPT-13B serving with random request arrivals at high load.
I want to be honest here: this number represents the best-case scenario. The improvement varies dramatically based on:
- Request rate: At low load, both systems handle everything fine. The gap only appears when memory becomes the bottleneck.
- Sequence length variance: vLLM shines when sequences have unpredictable lengths. If all your sequences are exactly 1000 tokens, the waste from pre-allocation is smaller.
- Batch size: Larger batches amplify the memory efficiency gains.
The paper’s Table 2 shows improvements ranging from 2.2x to 24.3x depending on the workload. For the ShareGPT dataset (real conversation traces), the improvement was about 3.5x for LLaMA-13B. Still significant, but not 24x.
What surprised me most in the ablations: the copy-on-write sharing contributes less than I expected for most workloads. The bulk of the improvement comes from eliminating fragmentation, not from beam search optimizations.

Implementation Details That Matter
If you’re planning to use vLLM in production, a few things worth knowing:
Block size selection: The default is 16 tokens per block. Smaller blocks mean less internal fragmentation but more pointer overhead. Larger blocks reduce metadata but waste more space. 16 is a reasonable default for models up to 30B parameters, but for larger models with bigger KV cache per token, you might want to tune this.
Preemption handling: When memory runs out, vLLM can preempt running sequences by swapping their KV cache blocks to CPU memory. This is a last resort—CPU↔GPU transfers are slow—but it prevents request failures. The swapping granularity matches the block size, so you can partially evict a sequence.
Continuous batching: vLLM batches requests at the iteration level, not the request level. New requests can join a batch mid-generation, and finished requests leave immediately. This is orthogonal to PagedAttention but critical for throughput.
Here’s a minimal example showing the difference in memory behavior:
# HuggingFace approach (simplified)
class NaiveKVCache:
def __init__(self, max_seq_len, batch_size, num_layers, num_heads, head_dim):
# Pre-allocate everything upfront
self.k_cache = torch.zeros(
(num_layers, batch_size, max_seq_len, num_heads, head_dim),
device='cuda'
) # This memory is reserved even for empty positions
self.v_cache = torch.zeros_like(self.k_cache)
# vLLM approach (conceptual, real impl is in CUDA)
class PagedKVCache:
def __init__(self, num_blocks, block_size, num_layers, num_heads, head_dim):
# Allocate a pool of blocks
self.block_pool = torch.zeros(
(num_blocks, num_layers, 2, block_size, num_heads, head_dim),
device='cuda'
) # 2 for K and V
self.free_blocks = list(range(num_blocks))
self.block_tables = {} # seq_id -> list of block indices
def allocate_block(self, seq_id):
if not self.free_blocks:
raise RuntimeError("OOM - need to preempt or reject")
block_idx = self.free_blocks.pop()
if seq_id not in self.block_tables:
self.block_tables[seq_id] = []
self.block_tables[seq_id].append(block_idx)
return block_idx
The real vLLM kernel is written in CUDA with careful attention to memory coalescing. My Python pseudocode above captures the concept but not the performance characteristics.
Comparison with Prior Work
PagedAttention didn’t emerge in a vacuum. A few relevant comparisons:
| System | Memory Management | Batching | Main Limitation |
|---|---|---|---|
| HuggingFace Transformers | Static pre-allocation | Request-level | Memory waste |
| FasterTransformer (NVIDIA) | Static pre-allocation | Iteration-level | Still wastes memory |
| Orca (OSDI ’22) | Dynamic allocation | Iteration-level | Fragmentation |
| vLLM | Paged blocks | Iteration-level | Kernel complexity |
The Orca system from Yu et al. (OSDI 2022) introduced continuous batching but didn’t solve the memory fragmentation problem. vLLM builds on Orca’s batching insights while adding the paging mechanism.
I covered attention mechanisms more broadly in LSTM Attention vs Self-Attention: How Bahdanau Evolved, but the optimization challenges at serving scale are a different beast entirely.
What the Paper Doesn’t Tell You
A few limitations worth noting:
Tensor parallelism overhead: When running across multiple GPUs, the block tables need to be synchronized. The paper doesn’t deeply analyze the cost of this coordination at large scale.
First-token latency: PagedAttention optimizes throughput, not latency. For real-time applications where time-to-first-token matters, the overhead of setting up block tables and allocation bookkeeping can be noticeable. My best guess is this adds 5-10ms for typical requests, though the paper doesn’t provide detailed latency breakdowns.
Model architecture assumptions: The paper focuses on decoder-only Transformers (GPT-style). Encoder-decoder models like T5 would need different handling for the cross-attention KV cache.
Quantization interaction: The paper doesn’t explore how PagedAttention interacts with KV cache quantization (INT8, FP8). In practice, combining paging with quantization might have non-obvious memory alignment issues.
For production inference cost optimization, I’ve found that combining memory-efficient attention with careful quantization gives better TCO than either approach alone—see my comparison in INT8 vs FP16 Inference: TCO Cut 54% for 7B Models on AWS.
Would I Use This in Production?
Absolutely, with caveats.
For high-throughput LLM serving (API endpoints, batch processing), vLLM with PagedAttention is the obvious choice over raw HuggingFace. The memory efficiency alone justifies the switch.
For latency-sensitive applications (real-time chat with strict SLAs), you might want to benchmark against TensorRT-LLM, which takes a different approach (highly optimized static graphs). The tradeoff isn’t clear-cut.
For research and experimentation where you’re frequently changing model architectures, HuggingFace’s flexibility might outweigh vLLM’s performance. vLLM requires models to be explicitly supported.
After staring at memory profiler output for longer than I’d like to admit (perhaps fueled by Dark Chocolate Espresso Beans), the thing that stands out is how much serving efficiency depends on workload characteristics. There’s no universal “best” system—just different tradeoffs for different deployment scenarios.
The Bigger Picture
PagedAttention represents a broader trend: applying classical systems techniques to ML infrastructure. Virtual memory paging is a 50-year-old idea. Copy-on-write fork semantics have been in Unix since the 1980s. What’s new is recognizing that these patterns apply to neural network serving.
I expect we’ll see more of this. Speculative execution for auto-regressive generation (already happening with speculative decoding). Cache coherence protocols for distributed inference. Garbage collection strategies for dynamic neural network graphs.
The question I haven’t figured out yet: as models get larger and context windows extend to millions of tokens, will block-based paging remain the right abstraction? At some point, the block table metadata itself becomes a memory burden. Maybe hierarchical page tables, like modern CPUs use for 64-bit address spaces?
But that’s a problem for next year’s papers.
FAQ
Q: Does PagedAttention work with any Transformer model?
vLLM supports most popular decoder-only architectures (LLaMA, GPT, Mistral, Falcon, etc.), but encoder-decoder models like T5 aren’t fully supported yet. The attention kernel needs model-specific adaptations, so check vLLM’s model compatibility list before assuming your model will work.
Q: How much memory overhead does the block table add?
The block table is tiny compared to the actual KV cache—typically a few KB per sequence for models with thousands of tokens. With 16-token blocks and 64-bit pointers, a 4096-token sequence needs only 256 block pointers (2KB). The overhead becomes noticeable only with millions of concurrent sequences.
Q: Can I combine PagedAttention with model quantization?
vLLM supports quantized models (AWQ, GPTQ, and others) alongside PagedAttention. The memory savings stack—you get both reduced KV cache precision and better memory utilization. Just be aware that block alignment requirements might change with different quantization schemes, potentially affecting how efficiently blocks pack.
References
- Kwon, W., Li, Z., Zhuang, S., et al. (2023). “Efficient Memory Management for Large Language Model Serving with PagedAttention.” SOSP 2023. https://arxiv.org/abs/2309.06180
- Yu, G. I., Jeong, J. S., Kim, G. W., et al. (2022). “Orca: A Distributed Serving System for Transformer-Based Generative Models.” OSDI 2022. https://www.usenix.org/conference/osdi22/presentation/yu
- Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). “Attention Is All You Need.” NeurIPS 2017. https://arxiv.org/abs/1706.03762
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,883 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (969 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (889 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (828 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (636 views)