PyTorch vs TensorFlow Syntax: 15 Operations Side-by-Side

Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links in this post are affiliate links — they cost you nothing extra.
⚡ Key Takeaways
  • PyTorch uses manual training loops with loss.backward() and optimizer.step(), while TensorFlow offers both model.fit() high-level API and GradientTape for custom loops
  • Key syntax differences: PyTorch uses torch.zeros(3, 4) vs TensorFlow's tf.zeros((3, 4)), and loss function argument order differs between frameworks
  • Both frameworks support mixed precision training (40% memory reduction), gradient clipping for RNN stability, and custom layers with similar subclassing patterns

Why Framework Cheat Sheets Beat Tutorials

You already know deep learning. You’ve built models in one framework, and now you need to read code in the other. You don’t need another 3000-word “gentle introduction” — you need the syntax for tensor slicing, custom loss functions, and checkpoint saving.

This is that reference. PyTorch 2.6 vs TensorFlow 2.18, covering the 15 operations you’ll actually use. Not a comprehensive guide (those exist), but the specific lines you’ll Google at 2am when translating someone else’s code.

I’ve been switching between both for three years — PyTorch for research experiments, TensorFlow for production deployment. The syntax differences aren’t huge, but they’re consistent enough to trip you up. Let’s fix that.

Visual abstraction of neural networks in AI technology, featuring data flow and algorithms.
Photo by Google DeepMind on Pexels

Tensor Basics: Creation and Indexing

PyTorch:

import torch

# Create tensors
x = torch.tensor([1, 2, 3], dtype=torch.float32)
y = torch.zeros(3, 4)  # shape (3, 4)
z = torch.randn(2, 5)  # normal distribution N(0,1)

# Indexing (zero-based, Python-style)
first_row = z[0, :]  # shape (5,)
subset = z[:, 1:3]   # columns 1-2, shape (2, 2)

TensorFlow:

import tensorflow as tf

# Create tensors
x = tf.constant([1, 2, 3], dtype=tf.float32)
y = tf.zeros((3, 4))  # note the tuple for shape
z = tf.random.normal((2, 5))

# Indexing (same syntax, but returns EagerTensor)
first_row = z[0, :]
subset = z[:, 1:3]

The big difference: PyTorch uses torch.zeros(3, 4) with positional args, TensorFlow uses tf.zeros((3, 4)) with a tuple. This trips up every migrator exactly once.

Enjoying this article? Get more like it delivered to your inbox. Subscribe to the newsletter

Device Management: CPU vs GPU

PyTorch:

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
x = torch.randn(100, 50).to(device)

# Or create directly on device
y = torch.randn(100, 50, device=device)

# Check current device
print(x.device)  # cuda:0 or cpu

TensorFlow:

# TensorFlow automatically places ops on GPU if available
# Manual placement only if needed:
with tf.device('/GPU:0'):
    x = tf.random.normal((100, 50))

# Check device
print(x.device)  # /job:localhost/replica:0/task:0/device:GPU:0

TensorFlow’s auto-placement is convenient until you hit multi-GPU setups. Then you’ll want explicit device scopes. PyTorch makes you think about placement from day one, which I actually prefer — fewer surprises in production.

Building a Simple Model

PyTorch (nn.Module):

import torch.nn as nn

class SimpleNet(nn.Module):
    def __init__(self, input_dim, hidden_dim, output_dim):
        super().__init__()
        self.fc1 = nn.Linear(input_dim, hidden_dim)
        self.relu = nn.ReLU()
        self.fc2 = nn.Linear(hidden_dim, output_dim)

    def forward(self, x):
        x = self.fc1(x)
        x = self.relu(x)
        x = self.fc2(x)
        return x

model = SimpleNet(784, 256, 10).to(device)

TensorFlow (Keras Sequential or Subclassing):

# Option 1: Sequential API (simpler)
model = tf.keras.Sequential([
    tf.keras.layers.Dense(256, activation='relu', input_shape=(784,)),
    tf.keras.layers.Dense(10)
])

# Option 2: Subclassing (PyTorch-style)
class SimpleNet(tf.keras.Model):
    def __init__(self, input_dim, hidden_dim, output_dim):
        super().__init__()
        self.fc1 = tf.keras.layers.Dense(hidden_dim, activation='relu')
        self.fc2 = tf.keras.layers.Dense(output_dim)

    def call(self, x):
        x = self.fc1(x)
        x = self.fc2(x)
        return x

model = SimpleNet(784, 256, 10)

PyTorch’s forward() vs TensorFlow’s call() — same idea, different names. Sequential API is faster to prototype, subclassing gives you full control for complex architectures.

Loss Functions

PyTorch:

criterion = nn.CrossEntropyLoss()  # expects raw logits
output = model(x)  # shape (batch, num_classes)
loss = criterion(output, labels)  # labels are class indices

# Custom loss (regression example)
def custom_loss(pred, target):
    mse = torch.mean((pred - target) ** 2)
    l1_penalty = 0.01 * torch.sum(torch.abs(pred))
    return mse + l1_penalty

TensorFlow:

loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)
output = model(x)
loss = loss_fn(labels, output)  # note: labels first in TensorFlow!

# Custom loss
def custom_loss(y_true, y_pred):
    mse = tf.reduce_mean(tf.square(y_pred - y_true))
    l1_penalty = 0.01 * tf.reduce_sum(tf.abs(y_pred))
    return mse + l1_penalty

Critical gotcha: PyTorch’s CrossEntropyLoss expects (predictions, labels), TensorFlow’s SparseCategoricalCrossentropy expects (labels, predictions). I’ve debugged this swap at least five times.

Optimizer Setup

PyTorch:

optimizer = torch.optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-5)

# Different learning rates for different layers
optimizer = torch.optim.Adam([
    {'params': model.fc1.parameters(), 'lr': 0.001},
    {'params': model.fc2.parameters(), 'lr': 0.0001}
])

TensorFlow:

optimizer = tf.keras.optimizers.Adam(learning_rate=0.001)

# Weight decay via optimizer (TF 2.x)
optimizer = tf.keras.optimizers.AdamW(learning_rate=0.001, weight_decay=1e-5)

# Layer-specific learning rates require custom training loop

PyTorch makes per-layer learning rates trivial. TensorFlow can do it, but you’ll need a custom training step. For 90% of cases, a single learning rate is fine.

Training Loop: The Core Difference

PyTorch (manual loop):

model.train()  # enable dropout/batchnorm training mode

for epoch in range(num_epochs):
    for batch_x, batch_y in train_loader:
        batch_x, batch_y = batch_x.to(device), batch_y.to(device)

        optimizer.zero_grad()  # CRITICAL: must clear gradients
        outputs = model(batch_x)
        loss = criterion(outputs, batch_y)
        loss.backward()  # compute gradients
        optimizer.step()  # update weights

TensorFlow (GradientTape for custom, or model.fit):

# Option 1: High-level API
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
model.fit(train_dataset, epochs=num_epochs, validation_data=val_dataset)

# Option 2: Custom training loop (PyTorch-style)
for epoch in range(num_epochs):
    for batch_x, batch_y in train_dataset:
        with tf.GradientTape() as tape:
            outputs = model(batch_x, training=True)
            loss = loss_fn(batch_y, outputs)

        gradients = tape.gradient(loss, model.trainable_variables)
        optimizer.apply_gradients(zip(gradients, model.trainable_variables))

This is where the frameworks diverge philosophically. PyTorch forces you to write the loop — you see every step, which I find clarifying. TensorFlow’s model.fit() is faster for standard cases but feels like magic until you peek under the hood with GradientTape.

And here’s the thing: if you’re doing anything non-standard (multi-task learning, adversarial training, curriculum learning), you’ll write a custom loop in both frameworks anyway.

Gradient Clipping

PyTorch:

loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()

TensorFlow:

with tf.GradientTape() as tape:
    loss = ...
gradients = tape.gradient(loss, model.trainable_variables)
clipped_grads, _ = tf.clip_by_global_norm(gradients, clip_norm=1.0)
optimizer.apply_gradients(zip(clipped_grads, model.trainable_variables))

Gradient clipping is mandatory for RNNs and Transformers. Without it, you’ll see loss spike to nan around epoch 3. The formula for global norm clipping is:

\mathbf{g} \leftarrow \frac{\text{clip_norm}}{\max(\text{clip_norm}, \|\mathbf{g}\|_2)} \cdot \mathbf{g}

where g\mathbf{g} is the concatenated gradient vector. Both frameworks use this.

Abstract 3D render visualizing artificial intelligence and neural networks in digital form.
Photo by Google DeepMind on Pexels

Custom Layers

PyTorch:

class ScaleLayer(nn.Module):
    def __init__(self, init_scale=1.0):
        super().__init__()
        self.scale = nn.Parameter(torch.tensor(init_scale))

    def forward(self, x):
        return x * self.scale

model = nn.Sequential(
    nn.Linear(128, 64),
    ScaleLayer(init_scale=0.5),
    nn.ReLU()
)

TensorFlow:

class ScaleLayer(tf.keras.layers.Layer):
    def __init__(self, init_scale=1.0):
        super().__init__()
        self.scale = self.add_weight(
            shape=(),
            initializer=tf.constant_initializer(init_scale),
            trainable=True
        )

    def call(self, x):
        return x * self.scale

model = tf.keras.Sequential([
    tf.keras.layers.Dense(64),
    ScaleLayer(init_scale=0.5),
    tf.keras.layers.ReLU()
])

PyTorch’s nn.Parameter vs TensorFlow’s add_weight() — same concept, different API. TensorFlow requires calling add_weight() in __init__, PyTorch lets you assign directly.

Saving and Loading Models

PyTorch:

# Save
torch.save({
    'epoch': epoch,
    'model_state_dict': model.state_dict(),
    'optimizer_state_dict': optimizer.state_dict(),
    'loss': loss
}, 'checkpoint.pth')

# Load
checkpoint = torch.load('checkpoint.pth')
model.load_state_dict(checkpoint['model_state_dict'])
optimizer.load_state_dict(checkpoint['optimizer_state_dict'])

TensorFlow:

# Save (Keras format)
model.save('model.keras')  # or model.save_weights('weights.h5')

# Load
model = tf.keras.models.load_model('model.keras')

# For checkpoints during training
checkpoint = tf.train.Checkpoint(optimizer=optimizer, model=model)
checkpoint.save('ckpt')
checkpoint.restore('ckpt-1')

PyTorch saves state dicts (weight tensors), TensorFlow saves the entire model architecture + weights by default. PyTorch’s approach gives you more control; TensorFlow’s is more convenient for quick experiments.

Data Loading

PyTorch (DataLoader):

from torch.utils.data import Dataset, DataLoader

class CustomDataset(Dataset):
    def __init__(self, data, labels):
        self.data = data
        self.labels = labels

    def __len__(self):
        return len(self.data)

    def __getitem__(self, idx):
        return self.data[idx], self.labels[idx]

train_dataset = CustomDataset(X_train, y_train)
train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True, num_workers=4)

for batch_x, batch_y in train_loader:
    # training code
    pass

TensorFlow (tf.data):

train_dataset = tf.data.Dataset.from_tensor_slices((X_train, y_train))
train_dataset = train_dataset.shuffle(buffer_size=1000).batch(32).prefetch(tf.data.AUTOTUNE)

for batch_x, batch_y in train_dataset:
    # training code
    pass

PyTorch’s DataLoader is more flexible for custom datasets. TensorFlow’s tf.data API is faster once you learn the chaining syntax (map, prefetch, cache). I’ve found tf.data wins for image augmentation pipelines, PyTorch wins for weird custom datasets.

Batch Normalization Gotcha

PyTorch:

model.train()  # batchnorm uses batch statistics
model.eval()   # batchnorm uses running mean/var

# During inference:
model.eval()
with torch.no_grad():
    outputs = model(x)

TensorFlow:

# Must pass training flag explicitly
outputs = model(x, training=True)  # training mode
outputs = model(x, training=False)  # inference mode

# Or use model.predict() which sets training=False
predictions = model.predict(x)

This bit me hard. In PyTorch, you toggle mode globally with model.train() / model.eval(). In TensorFlow, you pass training=True/False to the model call. Forgetting this during evaluation will give you slightly wrong results — not catastrophic, but enough to hurt your validation metrics.

Mixed Precision Training

PyTorch (AMP):

from torch.cuda.amp import autocast, GradScaler

scaler = GradScaler()

for batch_x, batch_y in train_loader:
    optimizer.zero_grad()

    with autocast():  # fp16 for forward pass
        outputs = model(batch_x)
        loss = criterion(outputs, batch_y)

    scaler.scale(loss).backward()  # scale gradients to prevent underflow
    scaler.step(optimizer)
    scaler.update()

TensorFlow:

from tensorflow.keras import mixed_precision
mixed_precision.set_global_policy('mixed_float16')

model = SimpleNet(784, 256, 10)
optimizer = tf.keras.optimizers.Adam()
optimizer = mixed_precision.LossScaleOptimizer(optimizer)

# Training loop handles scaling automatically
with tf.GradientTape() as tape:
    outputs = model(x, training=True)
    loss = loss_fn(y, outputs)
    scaled_loss = optimizer.get_scaled_loss(loss)

scaled_gradients = tape.gradient(scaled_loss, model.trainable_variables)
gradients = optimizer.get_unscaled_gradients(scaled_gradients)
optimizer.apply_gradients(zip(gradients, model.trainable_variables))

Mixed precision cuts memory usage by ~40% and speeds up training on modern GPUs (Ampere/Ada). The loss scaling L′=2kLL' = 2^k L (typically k=16k=16) prevents gradient underflow in fp16. Both frameworks implement this, syntax differs slightly.

Debugging: Checking for NaNs

PyTorch:

# Check for NaN in tensors
if torch.isnan(loss).any():
    print("NaN detected in loss!")

# Hook to catch NaN gradients
def check_nan_hook(grad):
    if grad is not None and torch.isnan(grad).any():
        raise ValueError("NaN gradient detected")

for param in model.parameters():
    param.register_hook(check_nan_hook)

TensorFlow:

# Check for NaN
if tf.reduce_any(tf.math.is_nan(loss)):
    print("NaN detected in loss!")

# Enable aggressive NaN checking (slow, debug only)
tf.debugging.enable_check_numerics()

NaN losses usually mean: (1) learning rate too high, (2) missing gradient clipping, (3) exploding gradients in recurrent layers, or (4) division by zero somewhere. The debugging tools above help you pinpoint where it starts.

Model Summary and Parameter Count

PyTorch:

from torchinfo import summary  # pip install torchinfo

summary(model, input_size=(32, 784))  # batch_size=32, input_dim=784

# Or manually count parameters
total_params = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Total trainable parameters: {total_params:,}")

TensorFlow:

model.summary()  # built-in, very clean output

# Or manually
model.count_params()

TensorFlow’s built-in model.summary() is excellent. PyTorch requires torchinfo (third-party), but it works well.

Transfer Learning: Freezing Layers

PyTorch:

# Freeze all layers
for param in model.parameters():
    param.requires_grad = False

# Unfreeze last layer
for param in model.fc2.parameters():
    param.requires_grad = True

# Only pass trainable params to optimizer
optimizer = torch.optim.Adam(filter(lambda p: p.requires_grad, model.parameters()), lr=0.001)

TensorFlow:

# Freeze entire model
model.trainable = False

# Freeze specific layer
model.layers[0].trainable = False

# Must recompile after changing trainable status
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy')

PyTorch’s requires_grad is more granular. TensorFlow requires recompiling the model after freezing layers, which always feels clunky to me.

Learning Rate Schedulers

PyTorch:

from torch.optim.lr_scheduler import ReduceLROnPlateau, CosineAnnealingLR

scheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=5)

# Or cosine annealing
scheduler = CosineAnnealingLR(optimizer, T_max=100, eta_min=1e-6)

# Call after each epoch
for epoch in range(num_epochs):
    train(...)
    val_loss = validate(...)
    scheduler.step(val_loss)  # for ReduceLROnPlateau

TensorFlow:

lr_schedule = tf.keras.optimizers.schedules.CosineDecay(
    initial_learning_rate=0.001,
    decay_steps=1000
)
optimizer = tf.keras.optimizers.Adam(learning_rate=lr_schedule)

# Or use callback for ReduceLROnPlateau
reduce_lr = tf.keras.callbacks.ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=5)
model.fit(train_dataset, epochs=num_epochs, callbacks=[reduce_lr])

The cosine annealing schedule follows ηt=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡(TcurTmax⁡π))\eta_t = \eta_{\min} + \frac{1}{2}(\eta_{\max} – \eta_{\min})(1 + \cos(\frac{T_{cur}}{T_{\max}}\pi)), which empirically improves final accuracy by 1-2% over constant learning rates in vision tasks.

Which One Should You Use?

Here’s my take after three years using both:

Use PyTorch if:
– You’re doing research or experimenting with novel architectures
– You want full control over the training loop
– You’re reading academic papers (90%+ use PyTorch in 2026)
– You value debugging transparency (.backward() + print() everywhere)

Use TensorFlow if:
– You’re deploying to production (TensorFlow Lite, TensorFlow Serving, TensorFlow.js)
– You want high-level APIs that handle boilerplate
– You’re working with TPUs (TensorFlow has better support)
– You’re building pipelines for mobile or edge devices

But honestly? Learn both. The syntax differences are superficial. The hard part — understanding backpropagation, regularization, optimization dynamics — transfers completely. If you know how ∂L∂W\frac{\partial L}{\partial W} flows through your network, the framework is just spelling.

What I’m still figuring out: how to make TensorFlow’s GradientTape feel as natural as PyTorch’s loss.backward(). The explicit tape context is conceptually cleaner, but in practice I find myself missing PyTorch’s brevity. Maybe it’s just muscle memory. Need a few more Dark Chocolate Espresso Beans and late-night sessions to rewire that.

FAQ

Q: Can I convert a PyTorch model to TensorFlow or vice versa?

Yes, but it’s painful. ONNX is the intermediate format — export from PyTorch to ONNX, then import into TensorFlow. Works for standard layers (conv, linear, pooling), breaks for custom ops. I’ve done this twice, wouldn’t recommend unless you have no choice.

Q: Which framework is faster for training?

Depends. PyTorch 2.x with torch.compile() matches TensorFlow’s XLA in most benchmarks (I covered this in PyTorch 2.6 vs TensorFlow 2.18: 5x Faster Training). For inference, TensorFlow Lite wins on mobile, PyTorch Mobile is catching up. On NVIDIA GPUs, they’re within 10% of each other.

Q: Should beginners start with PyTorch or TensorFlow?

PyTorch. The explicit training loop teaches you what’s actually happening. TensorFlow’s model.fit() is convenient but hides too much — you won’t understand why your model isn’t converging because you didn’t see the optimizer step. Learn PyTorch first, then TensorFlow’s high-level API makes sense as syntactic sugar.

The next thing I’m watching: JAX adoption. It’s NumPy + autograd + XLA, feels like PyTorch but compiles like TensorFlow. If it gets better ecosystem support (more pre-trained models, better docs), it could replace both for research. But that’s a 2027 conversation.

Did you find this helpful?

Your support keeps this blog running and ad-free content coming.

☕ Buy me a coffee
TODAY 1,800 | TOTAL 130,007