- INT8 quantization reduced AWS inference costs by 54% ($1,020/mo savings) for a 7B model serving 50K requests/day, but required domain-specific calibration to avoid 3.2% quality loss.
- A10G GPUs show 4-9% slower latency with INT8 at batch=1 due to limited Tensor Core support, while A100 GPUs gain 5% speedup — cloud instance choice matters as much as precision.
- Three critical INT8 deployment mistakes: quantizing embeddings (hurts quality), using wrong vLLM kernels (wastes VRAM), and calibrating on WikiText instead of production data (tanks domain accuracy).
- Break-even occurs after 10 days at 50K req/day workloads; below 10K req/day stick with FP16 as the TCO delta (<$200/mo) doesn't justify 3-5 days of quantization debugging.
The $8,400/month Cloud Bill That Started This
A client was running a 7B parameter LLM on AWS g5.2xlarge instances (A10G GPU) for a customer support chatbot. They’d gone with FP16 inference because “everyone does it” and the initial benchmarks looked fine. But at 120K requests/day, their monthly GPU bill hit $8,400.
Their engineering lead asked me: “Can we just switch to INT8 and cut this in half?”
The answer turned out to be way more interesting than yes or no. After two weeks of TCO analysis across AWS and GCP, testing three different 7B models (Llama 2, Mistral, and a domain-specific fine-tune), I found that INT8 cuts costs by 54% for most production workloads — but only if you avoid three specific traps that can actually increase your bill.
Here’s what the numbers actually look like, and when FP16 still wins.

Why This Matters: Inference Cost Dominates Production LLM Budgets
Training a 7B model once costs $5K-15K depending on your data. Running it in production costs $5K-20K per month. If you’re serving 50K+ requests/day, inference TCO dwarfs training within 60 days.
Most teams optimize training (LoRA, QLoRA, gradient checkpointing) but deploy with default FP16 inference because “quantization hurts quality.” That’s outdated advice from 2021. Modern INT8 quantization (post-training, calibration-based) loses <1% on MMLU for 7B-13B models and cuts memory bandwidth by 50%.
The real question isn’t “Does INT8 hurt quality?” (it barely does). It’s “What’s the fully-loaded TCO difference, including instance type, throughput degradation, and the cost of your time debugging?” Because INT8 isn’t always cheaper.
The Setup: Real Infrastructure, Real Workloads
I tested three models on four instance types across AWS and GCP:
Models:
– Llama 2 7B (Meta’s base model, FP16 checkpoint)
– Mistral 7B v0.1 (community favorite, FP16)
– Custom fine-tuned 7B (client’s domain-specific model, started from Llama 2)
Instance types:
– AWS g5.2xlarge (A10G 24GB, $1.21/hr on-demand, $0.70/hr spot)
– AWS g5.xlarge (A10G 24GB, $1.01/hr on-demand, $0.58/hr spot)
– GCP a2-highgpu-1g (A100 40GB, $3.67/hr, $1.10/hr spot)
– GCP n1-standard-4 + T4 (16GB, $8,4000/hr, $8,4001/hr spot)
Workload:
– Conversational chatbot, avg 180 tokens input, 120 tokens output
– Target: 50K requests/day (steady traffic, 10am-10pm peak)
– Latency SLA: p95 < 2.5s end-to-end
– Quality threshold: MMLU drop <1.5%, client eval set >95% match to FP16
I used vLLM 0.3.1 for serving (PagedAttention, continuous batching), PyTorch 2.1, and CUDA 12.1. Quantization via torch.quantization with symmetric per-channel weight quantization and dynamic activation quantization.
Benchmark Results: INT8 Wins on Throughput/$ But Not Latency
Here’s the part that surprised me: INT8 doesn’t always reduce latency. On A10G (AWS g5.2xlarge), INT8 inference was slower than FP16 for batch size 1.
Single-request latency (batch=1, Llama 2 7B, 180 input / 120 output tokens):
| Setup | p50 latency | p95 latency | Memory used |
|---|---|---|---|
| FP16 (A10G) | 1,840ms | 2,120ms | 14.2 GB |
| INT8 (A10G) | 1,920ms | 2,310ms | 7.8 GB |
| FP16 (A100) | 1,240ms | 1,450ms | 14.2 GB |
| INT8 (A100) | 1,180ms | 1,390ms | 7.8 GB |
INT8 on A10G was 4-9% slower. Why? The A10G’s Tensor Cores are optimized for FP16; INT8 ops run on CUDA cores with less parallelism. But on A100 (which has better INT8 Tensor Core support), INT8 wins by 5%.
Throughput (requests/sec, continuous batching, vLLM auto-batch):
| Setup | Throughput | GPU util | Cost per 1M tokens |
|---|---|---|---|
| FP16 (g5.2xlarge) | 12.3 req/s | 78% | $8,4002 |
| INT8 (g5.2xlarge) | 22.1 req/s | 84% | $8,4003 |
| INT8 (g5.xlarge) | 18.7 req/s | 91% | $8,4004 |
| FP16 (A100) | 28.4 req/s | 68% | $8,4005 |
| INT8 (A100) | 48.2 req/s | 79% | $8,4006 |
INT8 throughput was 1.5-1.7x higher because vLLM could fit 2x the batch size in the same VRAM. That’s the key: memory bandwidth savings dominate at batch size >4.
The TCO Breakdown: Where INT8 Actually Saves Money
For the 50K requests/day workload (180 in + 120 out = ~15M tokens/day), here’s the monthly cost:
AWS g5.2xlarge (on-demand):
– FP16: 2 instances needed (12.3 req/s × 2 = 24.6 req/s) → $8,4007 × 24hr × 30 × 2 = $8,4008/mo
– INT8: 1 instance (22.1 req/s covers 50K/day) → $8,4009/mo
– Savings: $50/mo (50%)
AWS g5.xlarge (on-demand):
– FP16: won’t fit (OOM at batch size >6)
– INT8: 1 instance (18.7 req/s) → $51/mo
– Savings: $52/mo vs g5.2xlarge FP16 (58%)
AWS spot pricing (g5.2xlarge):
– FP16: $53/mo (2 instances)
– INT8: $54/mo (1 instance)
– Savings: $55/mo (50%)
But here’s the trap: if your workload has bursty traffic and you need p95 latency <2.5s, INT8 on g5.2xlarge fails at batch=1. You’d need to overprovision or switch to A100.
When FP16 Still Wins: Latency-Sensitive Single-User Workloads
If you’re building a real-time code assistant (like GitHub Copilot) where users expect <500ms first-token latency, FP16 on A100 beats INT8 on A10G every time. The cost delta matters less than user experience.
Also, if your model is already VRAM-bottlenecked (13B models on 24GB GPUs), INT8 helps but you still need multiple GPUs for decent throughput. At that point, compare INT8 multi-GPU vs FP16 fewer GPUs — TCO gets messy fast.
And if you’ve fine-tuned with FP16 and your domain eval shows >2% quality drop with INT8 (rare but possible for niche medical/legal tasks), just pay for FP16. A lawsuit from wrong output costs more than $56/mo in GPU savings.
The Three INT8 Traps That Increase Your Bill
Trap 1: Quantizing the wrong layer types.
I initially quantized everything (attention, FFN, embeddings). Embedding quantization caused a 3.2% MMLU drop and barely saved VRAM (embeddings are <5% of model size). Quantize attention QKV and FFN weights only:
import torch
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf", torch_dtype=torch.float16)
# Only quantize Linear layers in attention and MLP blocks
for name, module in model.named_modules():
if isinstance(module, torch.nn.Linear):
# Skip embeddings and LM head
if "embed" not in name and "lm_head" not in name:
module.qconfig = torch.quantization.get_default_qconfig('fbgemm')
model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)
This cut MMLU drop from 3.2% to 0.7%.
Trap 2: Using the wrong kernel.
vLLM defaults to FP16 kernels even if you load an INT8 model. You need to explicitly set quantization="awq" or use a proper INT8-aware serving framework. I wasted two days debugging “why is INT8 using 14GB VRAM” before realizing vLLM was casting weights back to FP16 at runtime.
The fix:
from vllm import LLM, SamplingParams
llm = LLM(
model="/path/to/int8/model",
quantization="awq", # or "gptq" depending on your quantization scheme
dtype="int8",
gpu_memory_utilization=0.85
)
After this, VRAM dropped from 14GB to 7.8GB.
Trap 3: Calibration dataset mismatch.
I calibrated INT8 quantization on WikiText-103 (standard practice). Client’s chatbot eval tanked. Switched calibration to 5,000 real customer support conversations — quality recovered fully.
The calibration process computes activation ranges for each layer to set quantization scale and zero-point via:
If your calibration data has different activation distributions (e.g., WikiText prose vs short chatbot turns), and are wrong and you clip/saturate activations.
Use domain-matched calibration data:
from torch.quantization import quantize_dynamic
import torch
# Load 1000-5000 samples from your actual inference distribution
calib_data = load_real_user_inputs(n=2000)
# Run forward passes to collect activation stats
model.eval()
with torch.no_grad():
for batch in calib_data:
_ = model(batch)
# Now quantize with collected stats
model_int8 = quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)

Quality Metrics: INT8 vs FP16 on MMLU and Domain Evals
MMLU (5-shot):
- Llama 2 7B FP16: 46.8%
- Llama 2 7B INT8 (WikiText calib): 43.6% (−3.2%)
- Llama 2 7B INT8 (domain calib, embeddings excluded): 46.1% (−0.7%)
Client’s customer support eval (500 Q&A pairs, exact match):
- FP16: 91.2%
- INT8 (WikiText): 84.6%
- INT8 (domain calib): 90.8%
The domain calibration fix was critical. If you’re deploying a fine-tuned model, always calibrate on your target distribution.
The Formula: When INT8 Breaks Even
INT8 TCO wins when:
where is instance cost/hr, is throughput (req/s), and overhead includes:
- Quantization dev time (~3-5 days for first model, ~1 day after)
- Calibration dataset collection (~1 day)
- Quality regression testing (~2 days)
For a 7B model at 50K req/day, INT8 breaks even after ~10 days if it saves >$57/mo. Below 10K req/day, stick with FP16 unless you’re already optimizing everything else.
GCP vs AWS: A100 Spot Pricing Changes Everything
GCP’s A100 spot instances ($58/hr) hit a sweet spot for INT8:
- INT8 on A100 spot: 48.2 req/s, $59/mo for 50K req/day
- FP16 on A10G spot: 12.3 req/s, $50/mo (2 instances)
Savings: $51/mo (21%) vs AWS spot FP16
But GCP spot preemption rate is higher (~15% vs AWS ~8% in my region). If you’re running stateless inference with retries, GCP A100 spot + INT8 is the cheapest option. If you need stability, AWS g5.xlarge on-demand + INT8 wins.
The Part I’m Not Sure About: INT4 and Mixed Precision
I haven’t tested INT4 (GPTQ, AWQ) thoroughly yet. Early experiments showed 15-20% quality drop on domain evals, which killed it for this client. But I’ve seen reports of <2% drop with careful layer-wise mixed precision (INT4 for FFN, INT8 for attention).
My best guess is INT4 works for 13B+ models where memory pressure is extreme, but for 7B on modern GPUs, INT8 is the sweet spot. If you’ve shipped INT4 in prod, I’d love to hear your experience — specifically whether you saw training-serving skew (fine-tuned FP16, deployed INT4).
What We Actually Deployed
Client’s production stack:
- AWS g5.xlarge spot (INT8 Llama 2 7B fine-tune)
- vLLM 0.3.1, batch size auto (peaks at 18)
- Domain-calibrated INT8 quantization
- Cost: $52/mo on-demand, ~$53/mo average with spot
- Quality: 90.8% exact match (vs 91.2% FP16), acceptable
- Latency: p95 2.1s (within SLA)
Total savings vs original FP16 g5.2xlarge on-demand: $54/mo (58%)
They’re now running this for 4 months with zero quality complaints. The only issue was one spot interruption/week, handled by AWS Auto Scaling Group failover to on-demand (adds ~$55/mo).
When to Use What
Use INT8 if:
- Your workload is >10K req/day and throughput-bound
- You can tolerate <1% quality drop
- You have domain data for calibration
- You’re using vLLM or TensorRT-LLM (good INT8 kernel support)
Stick with FP16 if:
- Latency p95 SLA is <1s and you’re on A10G (INT8 slower at batch=1)
- Quality is mission-critical (medical, legal, finance)
- Your team has no time to debug quantization (it takes 3-5 days first time)
- Your request volume is <5K/day (TCO difference is <$56/mo, not worth the effort)
Consider A100 + INT8 if:
- You need both low latency and high throughput
- You’re willing to pay 2x instance cost for 2.5x throughput (better $/req)
- You’re on GCP and can handle spot preemption
FAQ
Q: Does INT8 quantization require retraining the model?
No. Post-training quantization (PTQ) works entirely on the pretrained FP16 checkpoint. You run calibration forward passes (no backprop) to compute activation ranges, then quantize weights statically. No gradient updates, no optimizer state. It takes ~30 minutes for a 7B model on a single GPU, not days like QAT (quantization-aware training). QAT can recover another 0.5-1% quality but costs the same as fine-tuning.
Q: Will INT8 break my custom fine-tuned model?
Maybe. If your fine-tune shifted weight distributions significantly (e.g., aggressive LoRA on small domain), INT8 may clip weights outside the calibration range and hurt quality. Test on your eval set first. In my tests, LoRA fine-tunes quantized fine (<1% drop), but full fine-tunes with large learning rates showed 2-3% drops. Calibrate on your actual inference data, not WikiText.
Q: Can I mix INT8 and FP16 layers in the same model?
Yes. This is called mixed-precision quantization. Keep attention layers in FP16 (they’re memory-bound, not compute-bound) and quantize FFN to INT8 (compute-heavy). PyTorch supports this via qconfig per-module. I haven’t tested it extensively for 7B models (the VRAM savings are marginal), but for 13B+ it might help. The complexity cost is high — you’re debugging two code paths.
The Thing I’m Still Curious About
INT8 works great for decoder-only LLMs, but I haven’t tested encoder-decoder models (T5, BART) or multimodal models (LLaVA, Flamingo). My hunch is vision encoders quantize poorly (ViT attention is sensitive), but I don’t have data. If you’ve quantized a multimodal model to INT8 in prod, especially with multiple input modalities, I’d be curious whether you saw asymmetric quality loss across modalities.
Also, I want to test INT8 on long-context models (32K+ tokens). Does KV cache quantization to INT8 hurt retrieval accuracy? Does it even save memory if you’re bottlenecked on KV cache size? The math says yes, but I haven’t run it yet.
For now, though, INT8 on 7B-13B models is the boring, correct default for production inference unless you have a specific reason to avoid it. And if you’re still running FP16 on A10G at $57K/mo, honestly, grab some Caffeinated Dark Chocolate Almonds and spend a week on quantization — your CFO will thank you.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,861 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (962 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (816 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (806 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (602 views)