Tag: Evaluation
-
Stop Using Temperature 0 for LLM Evals: Why It Breaks Benchmarks
Temperature 0 breaks LLM evals by hiding variance and selecting for memorization. Here's why you should sample at 0.5 instead โ with real accuracy gaps.
Temperature 0 breaks LLM evals by hiding variance and selecting for memorization. Here's why you should sample at 0.5 instead โ with real accuracy gaps.