Tag: model optimization
-
TFLite Model Conversion: 10 Commands That Actually Work
Ten TFLite conversion commands that actually work in production โ with the edge cases, quantization tradeoffs, and debugging tricks the docs skip.
-
TFLite & ONNX Mobile Setup: 2x Speed Left on Table
Default TFLite and ONNX configs waste 2-8x performance. Here's the exact setupโdelegates, threads, quantizationโthat closes the gap on ARM devices.
-
Speculative Decoding: Why 2x Faster Inference Fails
Speculative decoding promises 2x faster LLM inference, but real-world gains often disappoint. Debug the hidden bottlenecks killing your speedup.
-
TFLite vs ONNX Runtime: Pi Zero Latency at 32ms vs 89ms
Benchmark TFLite vs ONNX Runtime on Raspberry Pi Zero: which framework delivers faster inference? Latency comparison reveals clear winner.
-
ONNX Runtime Mobile: 8ms Inference on iPhone 13
Cut mobile inference from 200ms to 8ms by switching to ONNX Runtime. Benchmarks, CoreML quirks, and when TFLite still wins.
-
Whisper Architecture: How OpenAI’s Speech Model Works
Whisper's encoder-decoder uses 1.5GB VRAM for Large model. Break down the architecture and memory bottlenecks before mobile deployment.