Tag: INT8
-
ONNX INT8 vs FP16: 3x Latency Drop on Jetson Orin Nano
YOLOv8n INT8 cut latency from 47ms to 15ms on Jetson Orin Nano โ but small-object mAP dropped 5.7%. Real tradeoff numbers with power benchmarks.
-
INT8 vs FP16 Inference: TCO Cut 54% for 7B Models on AWS
INT8 quantization slashes AWS inference costs 54% vs FP16 for 7B LLMs. Real g5.xlarge benchmarks reveal the accuracy-speed-cost tradeoffs.
-
QAT vs PTQ: When 3% Accuracy Drop Kills Your Model
Compare QAT vs PTQ to find when that 3% accuracy gap destroys real-world performanceโand which quantization method saves your model.
-
Whisper Model Quantization for Mobile Deployment
Quantizing Whisper for mobile: post-training quantization failed, QAT crashed, but ONNX Runtime reduced model size 75% without accuracy loss.