Tag: LLM Inference
-
KV Cache Optimization: 3x Faster LLM Inference on 24GB VRAM
Learn KV cache optimization techniques to achieve 3x faster LLM inference with quantization, MQA, and PagedAttention on consumer GPUs with limited VRAM.
-
PagedAttention in vLLM: KV Cache Paging for 24x Throughput
vLLM's PagedAttention cuts KV cache waste from 60-80% to near zero. Real benchmarks show 2-24x throughput gains over HuggingFaceโhere's how paging works.
-
INT8 vs FP16 Inference: TCO Cut 54% for 7B Models on AWS
INT8 quantization slashes AWS inference costs 54% vs FP16 for 7B LLMs. Real g5.xlarge benchmarks reveal the accuracy-speed-cost tradeoffs.
-
Speculative Decoding: Why 2x Faster Inference Fails
Speculative decoding promises 2x faster LLM inference, but real-world gains often disappoint. Debug the hidden bottlenecks killing your speedup.
-
GQA Review: Grouped Query Attention for Faster LLM Inference
GQA reduces LLM inference memory 8x with only 0.1 ROUGE loss. How grouped KV heads and 5% uptraining work in Llama 2 and Mistral.
-
Speculative Decoding: How Medusa and EAGLE Speed Up LLMs
Medusa and EAGLE promise 2-3x LLM speedup via speculative decoding. Test results on LLaMA 2: acceptance rates, memory cost, and when it fails.
-
vLLM vs TensorRT-LLM: RTX 4090 Inference Benchmark
vLLM vs TensorRT-LLM head-to-head on RTX 4090 with Llama 3.1 8B โ throughput, latency, memory usage, and setup complexity compared. TensorRT-LLM wins 2.3x throughput but at what cost?