- PyTorch uses manual training loops with loss.backward() and optimizer.step(), while TensorFlow offers both model.fit() high-level API and GradientTape for custom loops
- Key syntax differences: PyTorch uses torch.zeros(3, 4) vs TensorFlow's tf.zeros((3, 4)), and loss function argument order differs between frameworks
- Both frameworks support mixed precision training (40% memory reduction), gradient clipping for RNN stability, and custom layers with similar subclassing patterns
Why Framework Cheat Sheets Beat Tutorials
You already know deep learning. You’ve built models in one framework, and now you need to read code in the other. You don’t need another 3000-word “gentle introduction” — you need the syntax for tensor slicing, custom loss functions, and checkpoint saving.
This is that reference. PyTorch 2.6 vs TensorFlow 2.18, covering the 15 operations you’ll actually use. Not a comprehensive guide (those exist), but the specific lines you’ll Google at 2am when translating someone else’s code.
I’ve been switching between both for three years — PyTorch for research experiments, TensorFlow for production deployment. The syntax differences aren’t huge, but they’re consistent enough to trip you up. Let’s fix that.

Tensor Basics: Creation and Indexing
PyTorch:
import torch
# Create tensors
x = torch.tensor([1, 2, 3], dtype=torch.float32)
y = torch.zeros(3, 4) # shape (3, 4)
z = torch.randn(2, 5) # normal distribution N(0,1)
# Indexing (zero-based, Python-style)
first_row = z[0, :] # shape (5,)
subset = z[:, 1:3] # columns 1-2, shape (2, 2)
TensorFlow:
import tensorflow as tf
# Create tensors
x = tf.constant([1, 2, 3], dtype=tf.float32)
y = tf.zeros((3, 4)) # note the tuple for shape
z = tf.random.normal((2, 5))
# Indexing (same syntax, but returns EagerTensor)
first_row = z[0, :]
subset = z[:, 1:3]
The big difference: PyTorch uses torch.zeros(3, 4) with positional args, TensorFlow uses tf.zeros((3, 4)) with a tuple. This trips up every migrator exactly once.
Device Management: CPU vs GPU
PyTorch:
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
x = torch.randn(100, 50).to(device)
# Or create directly on device
y = torch.randn(100, 50, device=device)
# Check current device
print(x.device) # cuda:0 or cpu
TensorFlow:
# TensorFlow automatically places ops on GPU if available
# Manual placement only if needed:
with tf.device('/GPU:0'):
x = tf.random.normal((100, 50))
# Check device
print(x.device) # /job:localhost/replica:0/task:0/device:GPU:0
TensorFlow’s auto-placement is convenient until you hit multi-GPU setups. Then you’ll want explicit device scopes. PyTorch makes you think about placement from day one, which I actually prefer — fewer surprises in production.
Building a Simple Model
PyTorch (nn.Module):
import torch.nn as nn
class SimpleNet(nn.Module):
def __init__(self, input_dim, hidden_dim, output_dim):
super().__init__()
self.fc1 = nn.Linear(input_dim, hidden_dim)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(hidden_dim, output_dim)
def forward(self, x):
x = self.fc1(x)
x = self.relu(x)
x = self.fc2(x)
return x
model = SimpleNet(784, 256, 10).to(device)
TensorFlow (Keras Sequential or Subclassing):
# Option 1: Sequential API (simpler)
model = tf.keras.Sequential([
tf.keras.layers.Dense(256, activation='relu', input_shape=(784,)),
tf.keras.layers.Dense(10)
])
# Option 2: Subclassing (PyTorch-style)
class SimpleNet(tf.keras.Model):
def __init__(self, input_dim, hidden_dim, output_dim):
super().__init__()
self.fc1 = tf.keras.layers.Dense(hidden_dim, activation='relu')
self.fc2 = tf.keras.layers.Dense(output_dim)
def call(self, x):
x = self.fc1(x)
x = self.fc2(x)
return x
model = SimpleNet(784, 256, 10)
PyTorch’s forward() vs TensorFlow’s call() — same idea, different names. Sequential API is faster to prototype, subclassing gives you full control for complex architectures.
Loss Functions
PyTorch:
criterion = nn.CrossEntropyLoss() # expects raw logits
output = model(x) # shape (batch, num_classes)
loss = criterion(output, labels) # labels are class indices
# Custom loss (regression example)
def custom_loss(pred, target):
mse = torch.mean((pred - target) ** 2)
l1_penalty = 0.01 * torch.sum(torch.abs(pred))
return mse + l1_penalty
TensorFlow:
loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)
output = model(x)
loss = loss_fn(labels, output) # note: labels first in TensorFlow!
# Custom loss
def custom_loss(y_true, y_pred):
mse = tf.reduce_mean(tf.square(y_pred - y_true))
l1_penalty = 0.01 * tf.reduce_sum(tf.abs(y_pred))
return mse + l1_penalty
Critical gotcha: PyTorch’s CrossEntropyLoss expects (predictions, labels), TensorFlow’s SparseCategoricalCrossentropy expects (labels, predictions). I’ve debugged this swap at least five times.
Optimizer Setup
PyTorch:
optimizer = torch.optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-5)
# Different learning rates for different layers
optimizer = torch.optim.Adam([
{'params': model.fc1.parameters(), 'lr': 0.001},
{'params': model.fc2.parameters(), 'lr': 0.0001}
])
TensorFlow:
optimizer = tf.keras.optimizers.Adam(learning_rate=0.001)
# Weight decay via optimizer (TF 2.x)
optimizer = tf.keras.optimizers.AdamW(learning_rate=0.001, weight_decay=1e-5)
# Layer-specific learning rates require custom training loop
PyTorch makes per-layer learning rates trivial. TensorFlow can do it, but you’ll need a custom training step. For 90% of cases, a single learning rate is fine.
Training Loop: The Core Difference
PyTorch (manual loop):
model.train() # enable dropout/batchnorm training mode
for epoch in range(num_epochs):
for batch_x, batch_y in train_loader:
batch_x, batch_y = batch_x.to(device), batch_y.to(device)
optimizer.zero_grad() # CRITICAL: must clear gradients
outputs = model(batch_x)
loss = criterion(outputs, batch_y)
loss.backward() # compute gradients
optimizer.step() # update weights
TensorFlow (GradientTape for custom, or model.fit):
# Option 1: High-level API
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
model.fit(train_dataset, epochs=num_epochs, validation_data=val_dataset)
# Option 2: Custom training loop (PyTorch-style)
for epoch in range(num_epochs):
for batch_x, batch_y in train_dataset:
with tf.GradientTape() as tape:
outputs = model(batch_x, training=True)
loss = loss_fn(batch_y, outputs)
gradients = tape.gradient(loss, model.trainable_variables)
optimizer.apply_gradients(zip(gradients, model.trainable_variables))
This is where the frameworks diverge philosophically. PyTorch forces you to write the loop — you see every step, which I find clarifying. TensorFlow’s model.fit() is faster for standard cases but feels like magic until you peek under the hood with GradientTape.
And here’s the thing: if you’re doing anything non-standard (multi-task learning, adversarial training, curriculum learning), you’ll write a custom loop in both frameworks anyway.
Gradient Clipping
PyTorch:
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
TensorFlow:
with tf.GradientTape() as tape:
loss = ...
gradients = tape.gradient(loss, model.trainable_variables)
clipped_grads, _ = tf.clip_by_global_norm(gradients, clip_norm=1.0)
optimizer.apply_gradients(zip(clipped_grads, model.trainable_variables))
Gradient clipping is mandatory for RNNs and Transformers. Without it, you’ll see loss spike to nan around epoch 3. The formula for global norm clipping is:
\mathbf{g} \leftarrow \frac{\text{clip_norm}}{\max(\text{clip_norm}, \|\mathbf{g}\|_2)} \cdot \mathbf{g}
where is the concatenated gradient vector. Both frameworks use this.

Custom Layers
PyTorch:
class ScaleLayer(nn.Module):
def __init__(self, init_scale=1.0):
super().__init__()
self.scale = nn.Parameter(torch.tensor(init_scale))
def forward(self, x):
return x * self.scale
model = nn.Sequential(
nn.Linear(128, 64),
ScaleLayer(init_scale=0.5),
nn.ReLU()
)
TensorFlow:
class ScaleLayer(tf.keras.layers.Layer):
def __init__(self, init_scale=1.0):
super().__init__()
self.scale = self.add_weight(
shape=(),
initializer=tf.constant_initializer(init_scale),
trainable=True
)
def call(self, x):
return x * self.scale
model = tf.keras.Sequential([
tf.keras.layers.Dense(64),
ScaleLayer(init_scale=0.5),
tf.keras.layers.ReLU()
])
PyTorch’s nn.Parameter vs TensorFlow’s add_weight() — same concept, different API. TensorFlow requires calling add_weight() in __init__, PyTorch lets you assign directly.
Saving and Loading Models
PyTorch:
# Save
torch.save({
'epoch': epoch,
'model_state_dict': model.state_dict(),
'optimizer_state_dict': optimizer.state_dict(),
'loss': loss
}, 'checkpoint.pth')
# Load
checkpoint = torch.load('checkpoint.pth')
model.load_state_dict(checkpoint['model_state_dict'])
optimizer.load_state_dict(checkpoint['optimizer_state_dict'])
TensorFlow:
# Save (Keras format)
model.save('model.keras') # or model.save_weights('weights.h5')
# Load
model = tf.keras.models.load_model('model.keras')
# For checkpoints during training
checkpoint = tf.train.Checkpoint(optimizer=optimizer, model=model)
checkpoint.save('ckpt')
checkpoint.restore('ckpt-1')
PyTorch saves state dicts (weight tensors), TensorFlow saves the entire model architecture + weights by default. PyTorch’s approach gives you more control; TensorFlow’s is more convenient for quick experiments.
Data Loading
PyTorch (DataLoader):
from torch.utils.data import Dataset, DataLoader
class CustomDataset(Dataset):
def __init__(self, data, labels):
self.data = data
self.labels = labels
def __len__(self):
return len(self.data)
def __getitem__(self, idx):
return self.data[idx], self.labels[idx]
train_dataset = CustomDataset(X_train, y_train)
train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True, num_workers=4)
for batch_x, batch_y in train_loader:
# training code
pass
TensorFlow (tf.data):
train_dataset = tf.data.Dataset.from_tensor_slices((X_train, y_train))
train_dataset = train_dataset.shuffle(buffer_size=1000).batch(32).prefetch(tf.data.AUTOTUNE)
for batch_x, batch_y in train_dataset:
# training code
pass
PyTorch’s DataLoader is more flexible for custom datasets. TensorFlow’s tf.data API is faster once you learn the chaining syntax (map, prefetch, cache). I’ve found tf.data wins for image augmentation pipelines, PyTorch wins for weird custom datasets.
Batch Normalization Gotcha
PyTorch:
model.train() # batchnorm uses batch statistics
model.eval() # batchnorm uses running mean/var
# During inference:
model.eval()
with torch.no_grad():
outputs = model(x)
TensorFlow:
# Must pass training flag explicitly
outputs = model(x, training=True) # training mode
outputs = model(x, training=False) # inference mode
# Or use model.predict() which sets training=False
predictions = model.predict(x)
This bit me hard. In PyTorch, you toggle mode globally with model.train() / model.eval(). In TensorFlow, you pass training=True/False to the model call. Forgetting this during evaluation will give you slightly wrong results — not catastrophic, but enough to hurt your validation metrics.
Mixed Precision Training
PyTorch (AMP):
from torch.cuda.amp import autocast, GradScaler
scaler = GradScaler()
for batch_x, batch_y in train_loader:
optimizer.zero_grad()
with autocast(): # fp16 for forward pass
outputs = model(batch_x)
loss = criterion(outputs, batch_y)
scaler.scale(loss).backward() # scale gradients to prevent underflow
scaler.step(optimizer)
scaler.update()
TensorFlow:
from tensorflow.keras import mixed_precision
mixed_precision.set_global_policy('mixed_float16')
model = SimpleNet(784, 256, 10)
optimizer = tf.keras.optimizers.Adam()
optimizer = mixed_precision.LossScaleOptimizer(optimizer)
# Training loop handles scaling automatically
with tf.GradientTape() as tape:
outputs = model(x, training=True)
loss = loss_fn(y, outputs)
scaled_loss = optimizer.get_scaled_loss(loss)
scaled_gradients = tape.gradient(scaled_loss, model.trainable_variables)
gradients = optimizer.get_unscaled_gradients(scaled_gradients)
optimizer.apply_gradients(zip(gradients, model.trainable_variables))
Mixed precision cuts memory usage by ~40% and speeds up training on modern GPUs (Ampere/Ada). The loss scaling (typically ) prevents gradient underflow in fp16. Both frameworks implement this, syntax differs slightly.
Debugging: Checking for NaNs
PyTorch:
# Check for NaN in tensors
if torch.isnan(loss).any():
print("NaN detected in loss!")
# Hook to catch NaN gradients
def check_nan_hook(grad):
if grad is not None and torch.isnan(grad).any():
raise ValueError("NaN gradient detected")
for param in model.parameters():
param.register_hook(check_nan_hook)
TensorFlow:
# Check for NaN
if tf.reduce_any(tf.math.is_nan(loss)):
print("NaN detected in loss!")
# Enable aggressive NaN checking (slow, debug only)
tf.debugging.enable_check_numerics()
NaN losses usually mean: (1) learning rate too high, (2) missing gradient clipping, (3) exploding gradients in recurrent layers, or (4) division by zero somewhere. The debugging tools above help you pinpoint where it starts.
Model Summary and Parameter Count
PyTorch:
from torchinfo import summary # pip install torchinfo
summary(model, input_size=(32, 784)) # batch_size=32, input_dim=784
# Or manually count parameters
total_params = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Total trainable parameters: {total_params:,}")
TensorFlow:
model.summary() # built-in, very clean output
# Or manually
model.count_params()
TensorFlow’s built-in model.summary() is excellent. PyTorch requires torchinfo (third-party), but it works well.
Transfer Learning: Freezing Layers
PyTorch:
# Freeze all layers
for param in model.parameters():
param.requires_grad = False
# Unfreeze last layer
for param in model.fc2.parameters():
param.requires_grad = True
# Only pass trainable params to optimizer
optimizer = torch.optim.Adam(filter(lambda p: p.requires_grad, model.parameters()), lr=0.001)
TensorFlow:
# Freeze entire model
model.trainable = False
# Freeze specific layer
model.layers[0].trainable = False
# Must recompile after changing trainable status
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy')
PyTorch’s requires_grad is more granular. TensorFlow requires recompiling the model after freezing layers, which always feels clunky to me.
Learning Rate Schedulers
PyTorch:
from torch.optim.lr_scheduler import ReduceLROnPlateau, CosineAnnealingLR
scheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=5)
# Or cosine annealing
scheduler = CosineAnnealingLR(optimizer, T_max=100, eta_min=1e-6)
# Call after each epoch
for epoch in range(num_epochs):
train(...)
val_loss = validate(...)
scheduler.step(val_loss) # for ReduceLROnPlateau
TensorFlow:
lr_schedule = tf.keras.optimizers.schedules.CosineDecay(
initial_learning_rate=0.001,
decay_steps=1000
)
optimizer = tf.keras.optimizers.Adam(learning_rate=lr_schedule)
# Or use callback for ReduceLROnPlateau
reduce_lr = tf.keras.callbacks.ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=5)
model.fit(train_dataset, epochs=num_epochs, callbacks=[reduce_lr])
The cosine annealing schedule follows , which empirically improves final accuracy by 1-2% over constant learning rates in vision tasks.
Which One Should You Use?
Here’s my take after three years using both:
Use PyTorch if:
– You’re doing research or experimenting with novel architectures
– You want full control over the training loop
– You’re reading academic papers (90%+ use PyTorch in 2026)
– You value debugging transparency (.backward() + print() everywhere)
Use TensorFlow if:
– You’re deploying to production (TensorFlow Lite, TensorFlow Serving, TensorFlow.js)
– You want high-level APIs that handle boilerplate
– You’re working with TPUs (TensorFlow has better support)
– You’re building pipelines for mobile or edge devices
But honestly? Learn both. The syntax differences are superficial. The hard part — understanding backpropagation, regularization, optimization dynamics — transfers completely. If you know how flows through your network, the framework is just spelling.
What I’m still figuring out: how to make TensorFlow’s GradientTape feel as natural as PyTorch’s loss.backward(). The explicit tape context is conceptually cleaner, but in practice I find myself missing PyTorch’s brevity. Maybe it’s just muscle memory. Need a few more Dark Chocolate Espresso Beans and late-night sessions to rewire that.
FAQ
Q: Can I convert a PyTorch model to TensorFlow or vice versa?
Yes, but it’s painful. ONNX is the intermediate format — export from PyTorch to ONNX, then import into TensorFlow. Works for standard layers (conv, linear, pooling), breaks for custom ops. I’ve done this twice, wouldn’t recommend unless you have no choice.
Q: Which framework is faster for training?
Depends. PyTorch 2.x with torch.compile() matches TensorFlow’s XLA in most benchmarks (I covered this in PyTorch 2.6 vs TensorFlow 2.18: 5x Faster Training). For inference, TensorFlow Lite wins on mobile, PyTorch Mobile is catching up. On NVIDIA GPUs, they’re within 10% of each other.
Q: Should beginners start with PyTorch or TensorFlow?
PyTorch. The explicit training loop teaches you what’s actually happening. TensorFlow’s model.fit() is convenient but hides too much — you won’t understand why your model isn’t converging because you didn’t see the optimizer step. Learn PyTorch first, then TensorFlow’s high-level API makes sense as syntactic sugar.
The next thing I’m watching: JAX adoption. It’s NumPy + autograd + XLA, feels like PyTorch but compiles like TensorFlow. If it gets better ecosystem support (more pre-trained models, better docs), it could replace both for research. But that’s a 2027 conversation.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,865 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (963 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (830 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (815 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (606 views)