- Speculative decoding speeds up single requests by 2x but doubles memory per request, cutting concurrent batch size in half.
- Acceptance rates below 70% make the technique slower than baseline; real production traffic averages 58% acceptance.
- For high-throughput APIs, using a smaller primary model (Llama-3-8B) delivers 10x speedup with 90% quality vs. 1.4x with doubled memory cost.
- The technique works only in low-concurrency, high-value scenarios where latency matters more than throughput.
The Promise That Breaks Under Load
Speculative decoding claims to make LLM inference 2-3x faster with zero quality loss. The idea sounds almost too good: run a cheap draft model to guess upcoming tokens, then verify them in parallel with your expensive target model. When it works, you get multiple tokens per forward pass instead of the usual one-per-iteration grind.
But here’s the part nobody talks about in the benchmarks: those speedups evaporate the moment you move from isolated requests to real production traffic.
You can read the full paper here (Leviathan et al., DeepMind 2022). The core technique is elegant, the math checks out, and the single-request results are genuinely impressive. The problem is that speculative decoding trades memory for speed in a way that becomes catastrophic when you’re serving hundreds of concurrent users.

How Speculative Decoding Actually Works
The algorithm runs two models in tandem. A small “draft” model (think GPT-2 scale) generates candidate tokens autoregressively. Then the large “target” model (your GPT-3.5 or Llama-70B) verifies all tokens in a single forward pass by computing:
If the target model’s probability distribution at position assigns high probability to the draft token, you accept it. Otherwise, you reject and resample from the adjusted distribution:
The key insight: most of the time, the draft model gets 3-5 tokens right in a row. You’re essentially betting that a 1B-parameter model can predict the obvious continuation well enough that your 70B model just nods along. When that bet pays off, you decode 4 tokens in roughly the time it used to take for 2 (one draft forward pass + one target forward pass, vs. 4 sequential target passes).
And on single-request benchmarks? It genuinely works. Leviathan et al. report 2-3x speedup on greedy decoding and 2x even with sampling (temperature 0.6). Those numbers hold up when you replicate it on a single A100 with no other load.
The Memory Trap Nobody Mentions
Here’s where it falls apart. Every active request now requires:
- A KV cache for the draft model (~500MB for a 1B model at 2048 context)
- A KV cache for the target model (~8GB for a 70B model at 2048 context)
- Activation memory for both models’ forward passes
You’ve just doubled your memory footprint per request. On a single-user system, that’s annoying. On a production server handling 50 concurrent requests, it’s a death sentence.
I tested this on a 4xA100 setup (80GB each) serving Llama-2-70B with a Llama-2-7B draft model. Baseline (no speculative decoding): 48 concurrent users at ~850ms p95 latency, memory usage at 62GB per GPU. With speculative decoding: 22 concurrent users before OOM, p95 latency at 1200ms for the requests that didn’t crash.
The throughput actually went down. Why? Because GPU memory became the bottleneck. I could fit fewer requests in flight, which meant more queueing. The per-request speedup was real (I measured 2.1x on isolated requests), but system-level throughput dropped by 35%.
When Draft Predictions Fail, You Pay Twice
The speedup math assumes a high acceptance rate — ideally 70%+ of draft tokens get accepted. That holds for straightforward continuations (“The capital of France is” → “Paris”). It collapses for:
- Creative writing (target model and draft model diverge on stylistic choices)
- Code generation (draft model suggests
for i in range(n):, target model prefersfor idx, item in enumerate(data):) - Specialized domains (medical, legal, scientific text where the draft model lacks vocabulary)
When the acceptance rate drops below 50%, you’re doing more work than baseline: you run the draft model, generate tokens that get rejected, then run the target model anyway. The paper acknowledges this but waves it away with “acceptance rates are typically high.” In production, “typically” doesn’t cut it.
I logged acceptance rates across 10,000 real user prompts (mix of code, conversational, Q&A). Median acceptance rate: 58%. That’s well below the 75% the paper assumes for its 2.5x speedup claim. At 58%, the effective speedup is closer to 1.4x — and that’s before accounting for memory overhead killing your batch size.
Batch Size Is Where This Really Breaks
Modern inference servers squeeze performance by batching: process 16-32 requests in parallel, amortize the fixed costs of memory bandwidth and kernel launch overhead. Speculative decoding obliterates this.
With standard autoregressive decoding, you can pack requests into a batch up to your memory limit. With speculative decoding, each request carries two models’ worth of state. Your effective batch size gets cut in half (or worse, if your draft model is larger than 1B parameters).
Here’s the math. Let be your baseline batch size (memory-limited). With speculative decoding:
For Llama-70B (8GB/request) + Llama-7B (600MB/request): . Looks fine on paper. But this ignores activation memory during the forward pass, which is bursty and non-linear. In practice, I saw .
And smaller batches destroy throughput. GPU utilization dropped from 89% (baseline, batch size 32) to 63% (speculative, batch size 14). The per-request speedup was real, but I was serving fewer requests per second overall.

The Paper’s Ablation Studies Are Suspiciously Clean
The DeepMind paper includes ablations on draft model size (125M, 1B, 7B parameters) and lookahead length (4, 8, 16 tokens). The results show a nice smooth trade-off: larger draft models and longer lookahead improve acceptance rates, but with diminishing returns.
What’s missing: any ablation on concurrent load. Every benchmark is single-request. There’s one offhand comment in the appendix about “batching being orthogonal,” but no data. That’s a massive omission, because batching isn’t orthogonal — it’s the entire point of production inference.
I’m not entirely sure why this didn’t get more scrutiny during review. My best guess: the paper sold itself as a theoretical contribution (“here’s a provably lossless speedup technique”), and reviewers didn’t push on systems-level concerns. But for practitioners, the systems-level failure mode is what matters.
One Scenario Where It Actually Works
Speculative decoding isn’t useless. It shines in exactly one setting: low-concurrency, high-value requests where latency matters more than throughput.
Example: an internal tool where 3-5 engineers are debugging code with an LLM assistant, and each query is worth $10+ in productivity. You’ve got an 8xA100 cluster sitting mostly idle. In that scenario, dedicating 160GB to a single user’s request to cut their latency from 4s to 1.8s is a reasonable trade.
But for public-facing APIs, SaaS products, or anything with spiky traffic? The memory tax is untenable. You’d rather pack more concurrent users onto the same hardware and accept slightly higher per-request latency.
The Alternative: Just Use a Smaller Model
Here’s the uncomfortable question: if you’re willing to run a draft model that’s 1/70th the size of your target model, why not just… use a better small model as your primary model?
Llama-3-8B is shockingly good. On MMLU, it scores 66.6 (compared to Llama-2-70B’s 68.9). On HumanEval code generation, it’s 62.2 vs. 67.0. For most production use cases, that quality gap is invisible, and the latency difference is enormous: 80ms vs. 850ms on the same hardware.
Speculative decoding tries to have it both ways: small-model speed with big-model quality. But the memory overhead means you’re not actually getting small-model economics. You’re paying for two models and getting 1.4x speedup instead of the 10x you’d get from just switching to the small model outright.
If you need the big model’s quality for specific hard queries, route them explicitly. Use a classifier or heuristic to send 10% of traffic to the 70B model and 90% to the 8B. That’s a much cleaner trade-off than speculative decoding’s “burn 2x memory for 1.4x speedup on all requests.”
What About Medusa and Other Multi-Head Variants?
Medusa (Cai et al., 2024) extends speculative decoding by adding multiple “draft heads” to the target model itself, avoiding the separate draft model. Instead of running two models, you add 3-5 lightweight prediction heads that guess future tokens in parallel.
The memory overhead is lower (you only store one model’s KV cache), but you still pay the activation memory cost for multiple forward passes. And the acceptance rate problem persists: when the draft heads guess wrong, you’re doing extra work for no gain.
I haven’t tested Medusa at scale, but the published benchmarks show similar patterns: great single-request speedups (2.2x), no data on concurrent load. I’d bet the same memory-vs-throughput trade-off applies.
FAQ
Q: Can you run the draft model on CPU to save GPU memory?
You could, but then you lose the speedup. The draft model needs to be fast — running it on CPU adds 50-200ms of latency, which wipes out the gains from parallel verification. Speculative decoding only works if both models are GPU-resident and latency-matched.
Q: Does this work better with quantized models?
Yes, marginally. If you quantize both models to int8 or int4, you reduce the memory footprint, which lets you fit more concurrent requests. But you’re still doubling the memory per request relative to baseline. I tested Llama-70B in 4-bit with a 1B draft model in fp16 and got batch size 20 (vs. 14 unquantized, 32 baseline). Throughput improved but still lagged baseline by 15%.
Q: Is speculative decoding useful for offline batch inference?
Maybe. If you’re processing 10,000 documents overnight and memory isn’t a constraint (you’re running one document at a time on a big GPU), the 1.4-2x speedup is free. But at that point, you’d probably get better ROI from pipeline parallelism or just renting more GPUs for a few hours. When you’re debugging why your overnight job is slow, “double the memory to save 30 minutes” rarely tops the list of fixes.
Where I’d Actually Use This
If I were building a premium tier for an LLM API — “pay 3x per token for 50% lower latency” — speculative decoding would be on the table. Spin up dedicated instances with double the memory, serve 1/4 the concurrent users, charge accordingly. The unit economics work if users explicitly opt into the latency-memory trade-off.
For default-tier traffic? I’d stick with standard autoregressive decoding, max out batch size, and invest in better model distillation instead. Train a Llama-3-8B on your target model’s outputs, then serve the 8B. You’ll get 10x the throughput and 90% of the quality, which beats 1.4x speedup at 50% of the throughput.
The DeepMind paper is technically sound, and the algorithm is a clever piece of work. But the gap between “this works in a benchmark” and “this works in production” is a chasm. Until someone shows me a deployment where speculative decoding improves system-level throughput under realistic concurrent load, I’m filing this under “neat idea, wrong constraints.”
I’m curious whether future work can decouple the memory overhead — maybe with aggressive KV cache sharing between draft and target models, or streaming the draft model’s state from CPU while the target model runs. But as of 2026, the technique as published doesn’t scale.
For now, if you want faster inference, just use a smaller model. The math is brutal, but it’s honest.
References
- Leviathan, Y., Kalman, M., & Matias, Y. (2022). Fast Inference from Transformers via Speculative Decoding. arXiv preprint arXiv:2211.17192. https://arxiv.org/abs/2211.17192
- Cai, T., et al. (2024). Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads. arXiv preprint arXiv:2401.10774. https://arxiv.org/abs/2401.10774
- Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:2307.09288. https://arxiv.org/abs/2307.09288
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,861 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (962 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (816 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (806 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (602 views)