Tag: Whisper
-
Whisper.cpp vs Faster-Whisper: Why Speed Tests Lie
Lab benchmarks show whisper.cpp winning, but production flips the winner. Here's why speed tests miss memory pressure, cold starts, and streaming overhead.
-
Whisper Tiny vs faster-whisper: 3x Speed, 12% WER Gap
faster-whisper cuts Jetson Nano latency 68% but WER jumps to 20%. Real benchmarks, memory spikes, and the thermal throttling nobody mentions.
-
Whisper Architecture: How OpenAI’s Speech Model Works
Whisper's encoder-decoder uses 1.5GB VRAM for Large model. Break down the architecture and memory bottlenecks before mobile deployment.
-
Real-time Whisper Is a Battery Nightmare (Here’s How to Fix It)
Real-time Whisper drains 1% battery/min. VAD + adaptive inference + thermal throttling bring it down to 0.2%. Benchmarks on iPhone 13 Pro.
-
On-Device Whisper: ONNX and Core ML Inference Guide
ONNX Runtime + CoreML beats native Core ML for Whisper on iOS by 40%. Conversion script, memory pool tricks, and A15/A16 benchmark numbers.
-
Whisper Model Quantization for Mobile Deployment
Quantizing Whisper for mobile: post-training quantization failed, QAT crashed, but ONNX Runtime reduced model size 75% without accuracy loss.
-
OpenAI Whisper ๋ชจ๋ธ ํ์ธํ๋ ์๋ฒฝ ๊ฐ์ด๋: ํ๊ตญ์ด ์์ฑ์ธ์ ์ ํ๋ ๋์ด๊ธฐ
Fine-tune OpenAI Whisper for Korean speech: boost accuracy from 82% to 94% with custom datasets. Includes training code and benchmarks.