- DPO doesn't eliminate the reward model — it defines an implicit reward as the log-ratio between policy and reference, scaled by β.
- RLHF uses ~40GB VRAM for a 7B model (4 models), while DPO needs ~20GB (2 models) or ~12GB with precomputed references.
- The β parameter controls deviation from the reference: too low causes reward hacking, too high prevents learning.
- DPO is offline-only; RLHF's online exploration advantage matters when preference data is sparse or limited.
- Length bias affects both methods differently — length normalization in DPO prevents verbose output gaming.
The Question That Stumped 80% of Candidates
“Walk me through how DPO eliminates the reward model.” Simple enough, right? I’ve sat through dozens of LLM interviews where candidates confidently explain that Direct Preference Optimization (DPO) is “simpler than RLHF” — then completely blank when asked to derive why. The math isn’t even that hard. The problem is that most tutorials hand you the final loss function without showing the sleight of hand that makes it work.
Here’s what actually trips people up: DPO doesn’t eliminate the reward model. It implicitly defines one. And that distinction matters when your interviewer asks follow-up questions.

RLHF’s Three-Stage Pipeline: Where the Complexity Lives
RLHF (Reinforcement Learning from Human Feedback), popularized by the InstructGPT paper (Ouyang et al., 2022), runs through three distinct phases. First, you supervised fine-tune (SFT) your base model on high-quality demonstrations. Second, you train a reward model on human preference pairs — given two responses, which is better? Third, you use PPO (Proximal Policy Optimization) to maximize that reward while staying close to your SFT checkpoint.
That third step is where things get messy.
PPO requires maintaining four models in memory during training: the policy being optimized, a frozen reference policy, the reward model, and a value function for variance reduction. On an A100 80GB, you’re looking at roughly 40GB just for a 7B parameter setup with all four models loaded. The optimization loop is notoriously unstable — reward hacking, KL divergence explosions, and mode collapse are constant companions. I’ve seen training runs where the reward kept climbing while actual response quality tanked because the model found a spurious pattern the reward model liked.
The PPO objective looks like this:
That KL penalty is doing a lot of heavy lifting. Too small and you get reward hacking. Too large and your model barely moves from the reference. Most teams I’ve talked to spend weeks just tuning this single hyperparameter.
DPO’s Trick: The Reward Is Already There
Direct Preference Optimization, introduced by Rafailov et al. (2023), starts from a clever observation. If you assume the reward function follows the Bradley-Terry model for pairwise preferences:
where is the sigmoid function, and you know that the optimal policy under KL-constrained reward maximization has the form:
then you can rearrange to express the reward in terms of the policy:
The partition function cancels out when you substitute back into the preference probability. This gives you the DPO loss directly:
No reward model. No PPO. No value function. Just standard cross-entropy-style optimization.
Interview Question 1: “What’s the Implicit Reward in DPO?”
This is where candidates stumble. The implicit reward is:
(ignoring the constant partition function)
But here’s what interviewers really want to hear: this reward is the log probability ratio between your trained policy and the reference, scaled by . Responses the model likes more than the reference get positive reward. Responses it likes less get negative reward. The model is essentially learning to predict its own preferences relative to its starting point.
The catch? This implicit reward can only represent preferences that are expressible as log-ratios. It’s less flexible than an arbitrary learned reward function. In practice this rarely matters, but it’s the kind of theoretical limitation that comes up in interviews.
Interview Question 2: “Why Does DPO Need a Reference Model?”
I’ve heard candidates say “DPO doesn’t need a reference model, that’s the whole point.” Wrong. DPO absolutely requires a reference policy , and you need to compute log probabilities from it during training.
The difference is:
– RLHF: reference model needed during PPO to compute KL penalty
– DPO: reference model needed to compute the log-ratio in the loss
Both need the reference. DPO just bakes it into the loss function rather than using it as a separate regularization term.
Here’s what this looks like in code:
import torch
import torch.nn.functional as F
def dpo_loss(
policy_chosen_logps: torch.Tensor, # log π_θ(y_w|x)
policy_rejected_logps: torch.Tensor, # log π_θ(y_l|x)
ref_chosen_logps: torch.Tensor, # log π_ref(y_w|x)
ref_rejected_logps: torch.Tensor, # log π_ref(y_l|x)
beta: float = 0.1,
) -> torch.Tensor:
"""
Compute DPO loss. This is surprisingly simple once you have the log probs.
"""
# The log-ratio terms
chosen_ratio = policy_chosen_logps - ref_chosen_logps
rejected_ratio = policy_rejected_logps - ref_rejected_logps
# Bradley-Terry preference model
logits = beta * (chosen_ratio - rejected_ratio)
# Binary cross-entropy where label is always 1 (chosen should be preferred)
loss = -F.logsigmoid(logits).mean()
# Some implementations track these for debugging
chosen_rewards = beta * chosen_ratio.detach()
rejected_rewards = beta * rejected_ratio.detach()
return loss, chosen_rewards, rejected_rewards
Notice the reference log probabilities. You can either load the reference model and compute these on-the-fly, or pre-compute them and store alongside your preference dataset. The latter is common for large-scale training — computing reference log probs once saves significant compute.
Interview Question 3: “When Does RLHF Still Win?”
DPO is simpler and often performs comparably. So why would anyone still use RLHF?
Online vs. offline learning. DPO is purely offline — it learns from a fixed preference dataset. RLHF with PPO can generate new samples during training, explore the response space, and potentially find better solutions than exist in your training data.
This matters when your preference data is limited or doesn’t cover the distribution you care about. If you have 10K preference pairs but want to deploy on a diverse set of prompts, RLHF’s ability to explore might help. DPO can only learn from what’s in the dataset.
There’s also the question of reward model interpretability. With RLHF, you can inspect your reward model, probe what it learned, and debug why certain responses score high. With DPO, the reward is implicit — you can extract it, but it’s harder to debug systematically.
My best guess is that for most practical applications, DPO wins on simplicity without sacrificing much quality. But if you’re doing cutting-edge alignment research where exploration matters, RLHF still has advantages.

Interview Question 4: “How Do You Handle Length Bias?”
Both methods suffer from it, but in different ways.
In RLHF, reward models notoriously favor longer responses. The model learns to pad outputs with hedging language, disclaimers, and repetition. You end up with responses that score high on the reward model but are frustrating to read.
DPO inherits this if your preference data has length bias. But it also has a more subtle issue: the log-probability ratio naturally correlates with length. Longer sequences have more tokens, and each token contributes to the log probability. A common fix is length normalization:
def length_normalized_dpo_loss(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
ref_chosen_logps: torch.Tensor,
ref_rejected_logps: torch.Tensor,
chosen_lengths: torch.Tensor,
rejected_lengths: torch.Tensor,
beta: float = 0.1,
) -> torch.Tensor:
"""
Length-normalized DPO. Dividing by sequence length prevents
the model from gaming the loss by generating longer outputs.
"""
# Average log prob per token
chosen_ratio = (policy_chosen_logps - ref_chosen_logps) / chosen_lengths
rejected_ratio = (policy_rejected_logps - ref_rejected_logps) / rejected_lengths
logits = beta * (chosen_ratio - rejected_ratio)
loss = -F.logsigmoid(logits).mean()
return loss
This isn’t in the original DPO paper but shows up in most production implementations. I’ve seen teams skip this and then wonder why their model suddenly becomes verbose after training.
Interview Question 5: “Explain the Parameter”
In both RLHF and DPO, controls how much the model can deviate from the reference policy.
- High (e.g., 0.5): strong regularization, model stays close to reference, conservative changes
- Low (e.g., 0.01): weak regularization, model can drift far, more aggressive optimization
The tricky part: the “right” value depends on your data quality and how different you want the final model to be. I’ve seen work well for general instruction tuning, but safety fine-tuning sometimes needs higher values to prevent catastrophic forgetting of safety behaviors.
Here’s a quick experiment showing how affects training dynamics:
import numpy as np
def simulate_dpo_gradient(beta_values, chosen_ratio, rejected_ratio):
"""
Show how beta affects gradient magnitude.
Larger beta means smaller effective gradients when ratios are similar.
"""
results = []
for beta in beta_values:
logit = beta * (chosen_ratio - rejected_ratio)
# Gradient of logsigmoid is (1 - sigmoid(x))
grad_scale = 1 - 1 / (1 + np.exp(-logit))
results.append((beta, logit, grad_scale))
return results
# Typical early training: chosen and rejected have similar log ratios
for beta, logit, grad in simulate_dpo_gradient([0.01, 0.1, 0.5], 0.5, 0.3):
print(f"beta={beta}: logit={logit:.3f}, grad_scale={grad:.3f}")
Output:
beta=0.01: logit=0.002, grad_scale=0.500
beta=0.1: logit=0.020, grad_scale=0.495
beta=0.5: logit=0.100, grad_scale=0.475
With low , gradients stay near 0.5 (maximum) even when the model is close to optimal. High causes gradients to drop faster as the model improves. This affects convergence speed and final solution quality in ways that aren’t always intuitive.
The Memory Comparison Interviewers Love
Peak GPU memory during training (7B parameter model):
| Method | Models in Memory | Approximate VRAM |
|---|---|---|
| RLHF (PPO) | Policy + Reference + Reward + Value | ~40GB |
| DPO | Policy + Reference | ~20GB |
| DPO (precomputed refs) | Policy only | ~12GB |
DPO’s memory advantage is real but often overstated. If you precompute reference log probabilities, you only need the policy model during training. But you still need to run inference on the reference model once per training example, so total compute isn’t always lower — it’s just shifted.
IPO and KTO: The DPO Variants You Should Know
Interviewers sometimes ask about DPO alternatives. Two worth knowing:
IPO (Identity Preference Optimization) from Azar et al. (2023) addresses a theoretical issue with DPO: as training progresses and the model becomes confident, DPO gradients vanish. IPO uses a different loss that maintains gradients:
KTO (Kahneman-Tversky Optimization) works with unpaired preference data — you just need “good” and “bad” examples, not explicit head-to-head comparisons. This is huge for practical data collection, since getting comparison pairs is expensive.
I haven’t tested KTO at scale myself, so take this with a grain of salt: early reports suggest it works nearly as well as DPO with much simpler data requirements.
Real Training Loop: Where Things Actually Break
Here’s a minimal but realistic DPO training loop. The comments highlight where I’ve seen things go wrong:
from transformers import AutoModelForCausalLM, AutoTokenizer
from torch.utils.data import DataLoader
import torch
def train_dpo(
model_name: str,
preference_data: list, # [{"prompt": ..., "chosen": ..., "rejected": ...}]
beta: float = 0.1,
epochs: int = 1,
lr: float = 1e-6, # DPO often needs very low LR
):
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load policy and reference (same init, but reference is frozen)
policy = AutoModelForCausalLM.from_pretrained(model_name).to(device)
ref_model = AutoModelForCausalLM.from_pretrained(model_name).to(device)
ref_model.eval()
for param in ref_model.parameters():
param.requires_grad = False
tokenizer = AutoTokenizer.from_pretrained(model_name)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token # Common gotcha
optimizer = torch.optim.AdamW(policy.parameters(), lr=lr)
for epoch in range(epochs):
for batch in preference_data:
# Tokenize chosen and rejected with prompt
chosen_ids = tokenizer(
batch["prompt"] + batch["chosen"],
return_tensors="pt",
truncation=True,
max_length=2048
).to(device)
rejected_ids = tokenizer(
batch["prompt"] + batch["rejected"],
return_tensors="pt",
truncation=True,
max_length=2048
).to(device)
# Get log probs (this is the tedious part)
with torch.no_grad():
ref_chosen_logps = get_sequence_logprob(ref_model, chosen_ids)
ref_rejected_logps = get_sequence_logprob(ref_model, rejected_ids)
policy_chosen_logps = get_sequence_logprob(policy, chosen_ids)
policy_rejected_logps = get_sequence_logprob(policy, rejected_ids)
loss, _, _ = dpo_loss(
policy_chosen_logps,
policy_rejected_logps,
ref_chosen_logps,
ref_rejected_logps,
beta=beta
)
optimizer.zero_grad()
loss.backward()
# Gradient clipping is essential — I've seen NaNs without it
torch.nn.utils.clip_grad_norm_(policy.parameters(), 1.0)
optimizer.step()
return policy
def get_sequence_logprob(model, inputs):
"""
Sum of log probs for all tokens in the response (excluding prompt).
This implementation is simplified — production code handles prompt masking.
"""
outputs = model(**inputs)
logits = outputs.logits[:, :-1, :] # Shift for next-token prediction
labels = inputs["input_ids"][:, 1:]
log_probs = torch.nn.functional.log_softmax(logits, dim=-1)
token_log_probs = log_probs.gather(-1, labels.unsqueeze(-1)).squeeze(-1)
# Mask padding tokens
mask = (labels != model.config.pad_token_id).float()
return (token_log_probs * mask).sum(dim=-1)
The get_sequence_logprob function is where most bugs hide. You need to properly handle prompt masking (only score response tokens), padding, and the off-by-one indexing between logits and labels. TRL library handles this correctly if you’d rather not debug it yourself.
What Most Tutorials Get Wrong
The biggest misconception I see: “DPO is just simpler RLHF.” It’s not. They solve different optimization problems that happen to converge to the same solution under specific assumptions.
RLHF maximizes expected reward with a KL constraint. DPO maximizes the likelihood that the model’s implicit preferences match human preferences. These are mathematically equivalent when the reward follows the Bradley-Terry model and you use the optimal policy form — but the training dynamics differ.
And sometimes DPO’s assumptions break down. If your preference data violates the Bradley-Terry assumption (e.g., preferences are intransitive), DPO can learn weird things. RLHF with a flexible reward model might handle this better, though I’m not entirely sure — I haven’t tested this systematically.
Debugging sessions at 2am are when you really appreciate the difference between these approaches. If you’re going to be up late staring at loss curves, at least have some dark chocolate espresso beans handy.
FAQ
Q: Can I use DPO without a reference model?
No. The reference model is mathematically required to compute the log-ratio that defines DPO’s implicit reward. Some implementations precompute reference log probabilities to avoid loading the model during training, but you still need it at some point. Removing the reference entirely would change the method fundamentally.
Q: Is DPO always better than RLHF for fine-tuning?
Not always. DPO is simpler and uses less memory, making it practical for smaller teams. But RLHF can explore beyond your training data through online sampling, which matters when your preference dataset is limited or doesn’t cover your deployment distribution. For most production use cases with decent preference data, DPO is the pragmatic choice.
Q: What value should I start with for DPO?
Start with for general instruction tuning. If your model drifts too far from the reference (incoherent outputs, forgotten capabilities), increase to 0.2-0.5. If it barely changes from the reference, try 0.05. Safety fine-tuning often needs higher values (0.3-0.5) to preserve existing safety behaviors.
Pick Your Method
Use DPO when you have good preference data and want straightforward training. It’s easier to debug, uses half the memory, and in most benchmarks performs comparably to RLHF. The TRL library from Hugging Face makes implementation trivial.
Use RLHF when you need online exploration, your preference data is sparse, or you want an interpretable reward model you can probe and debug. The complexity tax is real, but sometimes worth paying.
What I’m still curious about: how well do these methods scale beyond current model sizes? The implicit reward in DPO assumes a specific functional form — does that assumption hold as we push toward frontier models with emergent capabilities? The DPO paper showed strong results up to 7B parameters, but systematic comparisons at 70B+ are still scarce. That gap in the literature bothers me.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,861 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (962 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (815 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (806 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (602 views)