Category: LLM
-
vLLM vs TensorRT-LLM: RTX 4090 Inference Benchmark
vLLM vs TensorRT-LLM head-to-head on RTX 4090 with Llama 3.1 8B โ throughput, latency, memory usage, and setup complexity compared. TensorRT-LLM wins 2.3x throughput but at what cost?
-
Hybrid Search for RAG: BM25 + Vector Search Combined
Hybrid search boosts RAG recall 35% over vector-only. Combine BM25 + embeddings with rank fusionโPython code, real datasets, ablation results.
-
FlashAttention: Transformer ๋ฉ๋ชจ๋ฆฌ ์ต์ ํ ์ค์ ๊ตฌํ
FlashAttention-3 cuts Transformer memory from O(nยฒ) to O(n) on H100 GPUs. See PyTorch benchmarks proving 3-5x speedup on 32k context windows.
-
RoPE vs Alibi vs xPos: Transformer ์์น ์ธ์ฝ๋ฉ ์๋ฒฝ ๋น๊ต ๊ฐ์ด๋ (๊ธด ๋ฌธ๋งฅ ์ฒ๋ฆฌ ์ต์ ํ)
RoPE vs Alibi vs xPos on 128k token contexts: which Transformer position encoding prevents attention dilution? Benchmarks and PyTorch code inside.
-
Knowledge Distillation ์ค์ : LLM ์์ถ๊ณผ TinyLlama ์ฌ๋ก
Compress BERT to 40% size with 97% accuracy retained. Knowledge Distillation guide: DistilBERT and TinyLlama cases with PyTorch implementation.
-
RAG ํ์ดํ๋ผ์ธ ์ต์ ํ ์์ ๊ฐ์ด๋: Naive RAG๋ถํฐ Agentic RAG๊น์ง
Optimize RAG pipelines from naive to agentic: chunking strategies, hybrid search, and reranking code that actually improves retrieval.
-
Claude Code in Production: Best Practices and Automation
Claude Code CLI in production: quota fallback strategy, server-side KaTeX, and why it costs 94% less than API calls. Real numbers included.
-
Advanced Claude Code: Multi-Agent Workflows and Skills
Multi-agent pipelines in Claude Code solve the 200K token context wall. Here's how to chain agents without losing state or burning tokens.