Tag: vLLM
-
Ollama vs llama.cpp vs vLLM: Throughput on M1/RTX 4090
Compare Ollama, llama.cpp, and vLLM throughput on M1 and RTX 4090. Discover which framework delivers the best performance for local LLM inference.
-
Ollama vs vLLM vs llama.cpp: Which Wins for Your Use Case
vLLM hits 47x higher throughput than Ollama at 32 concurrent requests. Real benchmarks reveal when each framework wins โ and the memory tradeoffs nobody mentions.
-
KV Cache Optimization: 3x Faster LLM Inference on 24GB VRAM
Learn KV cache optimization techniques to achieve 3x faster LLM inference with quantization, MQA, and PagedAttention on consumer GPUs with limited VRAM.
-
PagedAttention in vLLM: KV Cache Paging for 24x Throughput
vLLM's PagedAttention cuts KV cache waste from 60-80% to near zero. Real benchmarks show 2-24x throughput gains over HuggingFaceโhere's how paging works.
-
vLLM OutOfMemoryError with Llama 3.1 70B: 3 Fixes
Fix vLLM OutOfMemoryError when deploying Llama 3.1 70B with tensor parallelism, quantization, and KV cache tuning on multi-GPU setups.
-
vLLM vs TensorRT-LLM: RTX 4090 Inference Benchmark
vLLM vs TensorRT-LLM head-to-head on RTX 4090 with Llama 3.1 8B โ throughput, latency, memory usage, and setup complexity compared. TensorRT-LLM wins 2.3x throughput but at what cost?