TFLite Model Conversion: 10 Commands That Actually Work

Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links in this post are affiliate links — they cost you nothing extra.
⚡ Key Takeaways
  • Dynamic range quantization cuts model size by 4x with <1% accuracy loss, but inference speed barely improves because activations stay float.
  • Full INT8 quantization requires a calibration dataset and drops inference time from 85ms to 38ms on Raspberry Pi 4 — mismatched calibration data causes silent accuracy drops.
  • LSTM models need hybrid quantization where recurrent layers stay FP32 while embeddings go INT8, achieving 3x size reduction with modest 20-30% speedup.

Why Most TFLite Conversion Examples Fail in Production

The TensorFlow Lite documentation shows you how to convert a model in three lines of code. What it doesn’t show: the seven edge cases that break silently in production, the quantization parameters that tank your accuracy, or why your converted model runs slower than the original.

I’ve converted dozens of models to TFLite for edge deployment — everything from MobileNet classifiers on Raspberry Pi to custom LSTM models on Android. The official examples work great for toy datasets. Real models? You’ll hit dtype mismatches, unsupported ops, and mysterious accuracy drops that take hours to debug.

Here are the ten commands I actually use, with the context you need to know when each one matters.

A digital representation of how large language models function in AI technology.
Photo by Google DeepMind on Pexels

The Baseline: SavedModel to TFLite (FP32)

Start here. If this doesn’t work, nothing else will.

import tensorflow as tf

converter = tf.lite.TFLiteConverter.from_saved_model('models/mobilenet_v2')
tflite_model = converter.convert()

with open('mobilenet_v2.tflite', 'wb') as f:
    f.write(tflite_model)

print(f"Model size: {len(tflite_model) / 1024 / 1024:.2f} MB")

This converts a SavedModel directory to TFLite format, keeping full 32-bit float precision. On a MobileNetV2 trained on ImageNet, you’ll get about 14 MB. Inference on a Raspberry Pi 4 takes around 85ms per image at 224×224 resolution.

But you’re not here for FP32. Edge devices need quantization.

Enjoying this article? Get more like it delivered to your inbox. Subscribe to the newsletter

Dynamic Range Quantization: The Easy 4x Size Win

This is the simplest quantization: weights go from FP32 to INT8, but activations stay float during inference. The math: weight tensor W∈Rm×nW \in \mathbb{R}^{m \times n} gets quantized as Wq=round(W−zs)W_q = \text{round}\left(\frac{W – z}{s}\right) where scale s=max⁡(W)−min⁡(W)255s = \frac{\max(W) – \min(W)}{255} and zero-point z=−round(min⁡(W)/s)z = -\text{round}(\min(W) / s).

converter = tf.lite.TFLiteConverter.from_saved_model('models/mobilenet_v2')
converter.optimizations = [tf.lite.Optimize.DEFAULT]
tflite_quant_model = converter.convert()

with open('mobilenet_v2_dynamic.tflite', 'wb') as f:
    f.write(tflite_quant_model)

print(f"Quantized size: {len(tflite_quant_model) / 1024 / 1024:.2f} MB")

MobileNetV2 shrinks from 14 MB to 3.5 MB. Accuracy drop on ImageNet is typically under 1% top-1. Inference time barely changes because the CPU still does float ops — you’re just loading less data from memory.

I use this when model size matters but I don’t have a representative dataset handy for full INT8 calibration.

Full Integer Quantization: The Real Speed Boost

This is where you get actual inference speedup on ARM CPUs and edge TPUs. Weights and activations go INT8. The catch: you need a representative dataset for calibration.

import numpy as np

def representative_dataset_gen():
    # Load 100-500 samples from your training set
    # Critical: these must match your real input distribution
    dataset = np.load('calibration_data.npy')  # Shape: (500, 224, 224, 3)
    for sample in dataset:
        # TFLite expects batch dimension
        yield [sample[np.newaxis, ...].astype(np.float32)]

converter = tf.lite.TFLiteConverter.from_saved_model('models/mobilenet_v2')
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset_gen
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8

tflite_int8_model = converter.convert()

On Raspberry Pi 4, this drops inference from 85ms to around 38ms for MobileNetV2. The model is still 3.5 MB (same as dynamic range), but now you’re using NEON SIMD integer instructions instead of floating point.

The representative_dataset_gen function is critical. I’ve seen people use random noise here — don’t. Pull actual samples from your training or validation set. Mismatched calibration data causes silent accuracy drops that you won’t catch until deployment.

Handling Unsupported Ops: The SELECT_TF_OPS Workaround

Some TensorFlow ops don’t have TFLite kernels. Custom layers, certain RNN cells, tf.image ops — the converter will error with “Some ops are not supported.”

converter = tf.lite.TFLiteConverter.from_saved_model('models/custom_model')
converter.target_spec.supported_ops = [
    tf.lite.OpsSet.TFLITE_BUILTINS,
    tf.lite.OpsSet.SELECT_TF_OPS  # Enable TensorFlow op fallback
]
converter.optimizations = [tf.lite.Optimize.DEFAULT]
tflite_model = converter.convert()

This bundles the full TensorFlow runtime ops into your TFLite model. The downside: model size balloons (add 1-2 MB of op kernels), and those ops won’t benefit from TFLite’s optimized kernels or quantization.

I use this as a temporary fix while I rewrite the unsupported layer. For deployment, you want pure TFLITE_BUILTINS.

Converting from Keras H5 Directly

If you have a .h5 file and no SavedModel, you can skip the intermediate step:

converter = tf.lite.TFLiteConverter.from_keras_model(
    tf.keras.models.load_model('model.h5')
)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
tflite_model = converter.convert()

This worked reliably in TensorFlow 2.8-2.12. In TF 2.13+, I’ve hit edge cases where custom layers with config dictionaries fail during conversion. If you control the training pipeline, save as SavedModel instead — it’s more robust.

Abstract representation of large language models and AI technology.
Photo by Google DeepMind on Pexels

Float16 Quantization: GPU Inference on Mobile

Float16 is the middle ground for GPU inference on mobile devices (Android, iOS). Weights are FP16, inference uses GPU half-precision ops.

converter = tf.lite.TFLiteConverter.from_saved_model('models/mobilenet_v2')
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.target_spec.supported_types = [tf.float16]
tflite_fp16_model = converter.convert()

Model size drops to ~7 MB (half of FP32). On a Pixel 6 with GPU acceleration enabled via the GPU delegate, inference is 15-20ms vs 45ms for FP32 CPU inference. Accuracy loss is negligible (under 0.1% for most vision models).

If you’re targeting iOS, CoreML might win here — I’ve seen cases where CoreML’s Metal backend beats TFLite’s GPU delegate by 30ms.

Batch Processing: Setting Input Shape at Conversion

By default, TFLite models accept dynamic batch sizes. If you know your batch size at deployment, lock it in:

converter = tf.lite.TFLiteConverter.from_saved_model('models/mobilenet_v2')
converter.optimizations = [tf.lite.Optimize.DEFAULT]
# Force batch size = 4
converter._experimental_lower_tensor_list_ops = False

tflite_model = converter.convert()

# Verify input shape
import tensorflow as tf
interpreter = tf.lite.Interpreter(model_content=tflite_model)
interpreter.allocate_tensors()
print(interpreter.get_input_details()[0]['shape'])  # [1, 224, 224, 3]

Actually, that example shows the default dynamic shape. To truly lock the batch size, you need to set it during model training or use tf.TensorSpec during SavedModel export. The converter itself doesn’t expose a clean batch-size-override API — this is one of those TensorFlow Lite rough edges.

In practice, most edge deployments use batch=1 anyway. If you need batching, you’re probably on a server and should reconsider whether TFLite is the right runtime.

Debugging Quantization Accuracy Loss

Your INT8 model dropped from 92% to 78% accuracy and you don’t know why. Here’s how to isolate the issue:

import tensorflow as tf
import numpy as np

# Load both models
interpreter_fp32 = tf.lite.Interpreter(model_path='model_fp32.tflite')
interpreter_int8 = tf.lite.Interpreter(model_path='model_int8.tflite')

interpreter_fp32.allocate_tensors()
interpreter_int8.allocate_tensors()

# Test on a single image
test_image = np.random.rand(1, 224, 224, 3).astype(np.float32)

interpreter_fp32.set_tensor(interpreter_fp32.get_input_details()[0]['index'], test_image)
interpreter_fp32.invoke()
output_fp32 = interpreter_fp32.get_tensor(interpreter_fp32.get_output_details()[0]['index'])

# For INT8, input needs scaling
input_details = interpreter_int8.get_input_details()[0]
scale, zero_point = input_details['quantization']
test_image_int8 = (test_image / scale + zero_point).astype(np.int8)

interpreter_int8.set_tensor(input_details['index'], test_image_int8)
interpreter_int8.invoke()
output_int8 = interpreter_int8.get_tensor(interpreter_int8.get_output_details()[0]['index'])

print(f"FP32 top-5: {np.argsort(output_fp32[0])[-5:][::-1]}")
print(f"INT8 top-5: {np.argsort(output_int8[0])[-5:][::-1]}")
print(f"Max output diff: {np.max(np.abs(output_fp32 - output_int8))}")

If the top-5 predictions diverge significantly, your calibration dataset is probably wrong. Go back and check that representative_dataset_gen actually covers the input distribution. I once debugged a 15% accuracy drop that turned out to be calibration images in BGR instead of RGB.

LSTM and RNN Quantization: The Hybrid Approach

Recurrent layers quantize poorly with full INT8 because hidden state activations have high dynamic range. The workaround: hybrid quantization, where RNN ops stay FP32 but everything else goes INT8.

converter = tf.lite.TFLiteConverter.from_saved_model('models/lstm_model')
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.target_spec.supported_ops = [
    tf.lite.OpsSet.TFLITE_BUILTINS,
    tf.lite.OpsSet.SELECT_TF_OPS
]
# Don't set inference_input_type/inference_output_type for hybrid mode
tflite_model = converter.convert()

This quantizes embeddings and dense layers but leaves LSTM cells in FP32. For a text classification model with a 128-unit LSTM, I got 3x size reduction (12 MB → 4 MB) with <1% accuracy loss. Inference speedup was modest (20-30%) because the LSTM still dominates runtime.

If you’re deploying audio or text models on edge, consider distilling the LSTM into a smaller Conv1D or Transformer encoder instead. Quantized convolutions run much faster than float LSTMs.

Edge TPU Compilation: The Final Step for Google Coral

If you’re targeting a Coral Edge TPU, you need an extra compilation step after TFLite conversion:

# Install Edge TPU compiler
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add -
echo "deb https://packages.cloud.google.com/apt coral-edgetpu-stable main" | sudo tee /etc/apt/sources.list.d/coral-edgetpu.list
sudo apt-get update
sudo apt-get install edgetpu-compiler

# Compile INT8 TFLite model
edgetpu_compiler mobilenet_v2_int8.tflite

This produces mobilenet_v2_int8_edgetpu.tflite. The compiler maps ops to the Edge TPU’s hardware and splits unsupported ops to CPU. Check the compiler output — if more than 20% of ops fall back to CPU, your model won’t benefit much from the TPU.

MobileNetV2 gets ~95% ops on TPU and runs in 3-5ms on a Coral USB Accelerator. Custom models with unsupported ops (some pooling layers, certain activation functions) see worse partitioning. Need more speed? Grab a Coral USB Accelerator and watch your latency drop by 10x.

What About PyTorch Models?

TFLite is a TensorFlow runtime. If your model is in PyTorch, export to ONNX first, then convert ONNX to TFLite via onnx-tf:

pip install onnx onnx-tf

# PyTorch → ONNX
torch.onnx.export(model, dummy_input, "model.onnx", opset_version=13)

# ONNX → TensorFlow SavedModel
onnx-tf convert -i model.onnx -o models/saved_model

# SavedModel → TFLite
python -c "
import tensorflow as tf
converter = tf.lite.TFLiteConverter.from_saved_model('models/saved_model')
converter.optimizations = [tf.lite.Optimize.DEFAULT]
open('model.tflite', 'wb').write(converter.convert())
"

This pipeline is brittle. ONNX → TensorFlow conversion fails on custom ops, dynamic shapes, and certain control flow patterns. I’ve had better luck with ONNX Runtime Mobile for PyTorch models — the toolchain is cleaner and you skip the ONNX→TF→TFLite dance.

When Conversion Still Fails

Some models just won’t convert cleanly. I’ve hit this with:
– Models using tf.py_function or tf.numpy_function (these can’t serialize)
– Custom layers that don’t implement get_config() properly
– Models with dynamic control flow (tf.cond, tf.while_loop) that TFLite can’t unroll

The nuclear option: rewrite the problematic layer in pure Keras/TF ops. I once spent four hours rewriting a custom attention layer to avoid a tf.py_function call. It worked.

Alternatively, if you control the training script, redesign the model architecture to use only TFLite-compatible ops from the start. Check the TensorFlow Lite op compatibility guide before training.

FAQ

Q: Why is my INT8 model slower than FP32 on Android?

You probably didn’t enable the NNAPI or GPU delegate. TFLite defaults to CPU inference, and without delegates, INT8 ops might run on a slow fallback kernel. Add interpreter.modifyGraphWithDelegate(NnApiDelegate()) in your Android code to use hardware acceleration. Also check that your SoC actually has INT8 acceleration — older Snapdragon 600-series chips don’t.

Q: Can I quantize just specific layers instead of the whole model?

Not directly via the converter API. You’d need to manually split the model, convert each part separately, and stitch them together — which defeats the purpose of end-to-end optimization. If certain layers are sensitive to quantization (e.g., first conv layer), consider using quantization-aware training (QAT) instead of post-training quantization. QAT inserts fake quantization nodes during training so the model learns to compensate for quantization error.

Q: What’s the difference between TFLite and TensorFlow.js?

TFLite targets native mobile/embedded (C++ runtime, optimized for ARM/EdgeTPU). TensorFlow.js runs in browsers via WebGL/WASM. For web deployment, use TF.js. For native apps or microcontrollers, use TFLite. TF.js models are typically 10-50% slower than equivalent TFLite models on the same device due to browser overhead, but you skip the app store and get instant updates.

Picking the Right Quantization Strategy

Here’s what I actually use:

  • Prototyping / model size doesn’t matter: FP32 baseline, no quantization
  • Model size matters, speed doesn’t: Dynamic range quantization (Optimize.DEFAULT only)
  • Inference speed matters, I have calibration data: Full INT8 quantization with representative_dataset
  • GPU inference on flagship Android/iOS: Float16 quantization
  • Coral Edge TPU deployment: Full INT8 + Edge TPU compiler
  • LSTM/RNN models: Hybrid quantization (let RNN ops stay FP32)

The hard part isn’t running the converter — it’s validating that your quantized model still works in production. Always measure accuracy on a held-out test set before deploying. I’ve seen 5% accuracy drops that were acceptable in context, and 0.5% drops that broke product requirements. Know your tolerance.

One thing I haven’t figured out: predicting quantization impact without running the full conversion. Post-training quantization is fast (seconds to minutes), but for giant models (BERT-large, ResNet-152), even calibration takes hours. Some papers propose sensitivity analysis to predict per-layer quantization impact, but I haven’t seen production tools for this yet. If you’ve cracked this, I’d love to hear how.

Did you find this helpful?

Your support keeps this blog running and ad-free content coming.

☕ Buy me a coffee
TODAY 1,385 | TOTAL 129,592