Tag: Model Quantization
-
TFLite Model Conversion: 10 Commands That Actually Work
Ten TFLite conversion commands that actually work in production โ with the edge cases, quantization tradeoffs, and debugging tricks the docs skip.
-
ONNX INT8 vs FP16: 3x Latency Drop on Jetson Orin Nano
YOLOv8n INT8 cut latency from 47ms to 15ms on Jetson Orin Nano โ but small-object mAP dropped 5.7%. Real tradeoff numbers with power benchmarks.
-
INT8 vs FP16 Inference: TCO Cut 54% for 7B Models on AWS
INT8 quantization slashes AWS inference costs 54% vs FP16 for 7B LLMs. Real g5.xlarge benchmarks reveal the accuracy-speed-cost tradeoffs.
-
Whisper Model Quantization for Mobile Deployment
Quantizing Whisper for mobile: post-training quantization failed, QAT crashed, but ONNX Runtime reduced model size 75% without accuracy loss.