Tag: INT8 quantization
-
Pruned YOLOv8 ONNX INT8 Fails: 3 Fixes That Work
Pruned YOLOv8 + ONNX INT8 dtype mismatch? Here are 3 working fixes with Jetson benchmarks โ re-quantize, QAT pruning, or manual graph surgery.
-
INT8 vs INT4 Quantization: 2x Latency Drop on ARM Cortex-M
INT4 quantization cuts Cortex-M inference latency in half โ but costs 18KB flash, breaks on residual nets, and drops accuracy 4-6% on edge cases.
-
YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin
INT8 quantization pushed YOLOv8 from 45 to 180 FPS on Jetson Orin Nano. PTQ beat QAT with 10 lines of calibration code. Here's the recipe.