DPO vs RLHF: 5 Interview Questions That Trip Up Developers

Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links in this post are affiliate links — they cost you nothing extra.
⚡ Key Takeaways
  • DPO doesn't eliminate the reward model — it defines an implicit reward as the log-ratio between policy and reference, scaled by β.
  • RLHF uses ~40GB VRAM for a 7B model (4 models), while DPO needs ~20GB (2 models) or ~12GB with precomputed references.
  • The β parameter controls deviation from the reference: too low causes reward hacking, too high prevents learning.
  • DPO is offline-only; RLHF's online exploration advantage matters when preference data is sparse or limited.
  • Length bias affects both methods differently — length normalization in DPO prevents verbose output gaming.

The Question That Stumped 80% of Candidates

“Walk me through how DPO eliminates the reward model.” Simple enough, right? I’ve sat through dozens of LLM interviews where candidates confidently explain that Direct Preference Optimization (DPO) is “simpler than RLHF” — then completely blank when asked to derive why. The math isn’t even that hard. The problem is that most tutorials hand you the final loss function without showing the sleight of hand that makes it work.

Here’s what actually trips people up: DPO doesn’t eliminate the reward model. It implicitly defines one. And that distinction matters when your interviewer asks follow-up questions.

Wooden Scrabble tiles spelling 'Deepmind' and 'Gemini' on a wooden surface, a concept of AI and games.
Photo by Markus Winkler on Pexels

RLHF’s Three-Stage Pipeline: Where the Complexity Lives

RLHF (Reinforcement Learning from Human Feedback), popularized by the InstructGPT paper (Ouyang et al., 2022), runs through three distinct phases. First, you supervised fine-tune (SFT) your base model on high-quality demonstrations. Second, you train a reward model on human preference pairs — given two responses, which is better? Third, you use PPO (Proximal Policy Optimization) to maximize that reward while staying close to your SFT checkpoint.

That third step is where things get messy.

PPO requires maintaining four models in memory during training: the policy being optimized, a frozen reference policy, the reward model, and a value function for variance reduction. On an A100 80GB, you’re looking at roughly 40GB just for a 7B parameter setup with all four models loaded. The optimization loop is notoriously unstable — reward hacking, KL divergence explosions, and mode collapse are constant companions. I’ve seen training runs where the reward kept climbing while actual response quality tanked because the model found a spurious pattern the reward model liked.

The PPO objective looks like this:

LPPO=Ex∼D,y∼πθ[rϕ(x,y)−β⋅DKL(πθ(y∣x)∥πref(y∣x))]\mathcal{L}_{PPO} = \mathbb{E}_{x \sim D, y \sim \pi_\theta}\left[r_\phi(x, y) – \beta \cdot D_{KL}(\pi_\theta(y|x) \| \pi_{ref}(y|x))\right]

That KL penalty β\beta is doing a lot of heavy lifting. Too small and you get reward hacking. Too large and your model barely moves from the reference. Most teams I’ve talked to spend weeks just tuning this single hyperparameter.

Enjoying this article? Get more like it delivered to your inbox. Subscribe to the newsletter

DPO’s Trick: The Reward Is Already There

Direct Preference Optimization, introduced by Rafailov et al. (2023), starts from a clever observation. If you assume the reward function follows the Bradley-Terry model for pairwise preferences:

p(yw≻yl∣x)=σ(r(x,yw)−r(x,yl))p(y_w \succ y_l | x) = \sigma(r(x, y_w) – r(x, y_l))

where σ\sigma is the sigmoid function, and you know that the optimal policy under KL-constrained reward maximization has the form:

π∗(y∣x)=1Z(x)πref(y∣x)exp⁡(r(x,y)β)\pi^*(y|x) = \frac{1}{Z(x)} \pi_{ref}(y|x) \exp\left(\frac{r(x,y)}{\beta}\right)

then you can rearrange to express the reward in terms of the policy:

r(x,y)=βlog⁡π∗(y∣x)πref(y∣x)+βlog⁡Z(x)r(x, y) = \beta \log \frac{\pi^*(y|x)}{\pi_{ref}(y|x)} + \beta \log Z(x)

The partition function Z(x)Z(x) cancels out when you substitute back into the preference probability. This gives you the DPO loss directly:

LDPO=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{DPO} = -\mathbb{E}_{(x, y_w, y_l) \sim D}\left[\log \sigma\left(\beta \log \frac{\pi_\theta(y_w|x)}{\pi_{ref}(y_w|x)} – \beta \log \frac{\pi_\theta(y_l|x)}{\pi_{ref}(y_l|x)}\right)\right]

No reward model. No PPO. No value function. Just standard cross-entropy-style optimization.

Interview Question 1: “What’s the Implicit Reward in DPO?”

This is where candidates stumble. The implicit reward is:

r(x,y)=βlog⁡πθ(y∣x)πref(y∣x)r(x, y) = \beta \log \frac{\pi_\theta(y|x)}{\pi_{ref}(y|x)}

(ignoring the constant partition function)

But here’s what interviewers really want to hear: this reward is the log probability ratio between your trained policy and the reference, scaled by β\beta. Responses the model likes more than the reference get positive reward. Responses it likes less get negative reward. The model is essentially learning to predict its own preferences relative to its starting point.

The catch? This implicit reward can only represent preferences that are expressible as log-ratios. It’s less flexible than an arbitrary learned reward function. In practice this rarely matters, but it’s the kind of theoretical limitation that comes up in interviews.

Interview Question 2: “Why Does DPO Need a Reference Model?”

I’ve heard candidates say “DPO doesn’t need a reference model, that’s the whole point.” Wrong. DPO absolutely requires a reference policy πref\pi_{ref}, and you need to compute log probabilities from it during training.

The difference is:
– RLHF: reference model needed during PPO to compute KL penalty
– DPO: reference model needed to compute the log-ratio in the loss

Both need the reference. DPO just bakes it into the loss function rather than using it as a separate regularization term.

Here’s what this looks like in code:

import torch
import torch.nn.functional as F

def dpo_loss(
    policy_chosen_logps: torch.Tensor,    # log π_θ(y_w|x)
    policy_rejected_logps: torch.Tensor,  # log π_θ(y_l|x)
    ref_chosen_logps: torch.Tensor,       # log π_ref(y_w|x)
    ref_rejected_logps: torch.Tensor,     # log π_ref(y_l|x)
    beta: float = 0.1,
) -> torch.Tensor:
    """
    Compute DPO loss. This is surprisingly simple once you have the log probs.
    """
    # The log-ratio terms
    chosen_ratio = policy_chosen_logps - ref_chosen_logps
    rejected_ratio = policy_rejected_logps - ref_rejected_logps

    # Bradley-Terry preference model
    logits = beta * (chosen_ratio - rejected_ratio)

    # Binary cross-entropy where label is always 1 (chosen should be preferred)
    loss = -F.logsigmoid(logits).mean()

    # Some implementations track these for debugging
    chosen_rewards = beta * chosen_ratio.detach()
    rejected_rewards = beta * rejected_ratio.detach()

    return loss, chosen_rewards, rejected_rewards

Notice the reference log probabilities. You can either load the reference model and compute these on-the-fly, or pre-compute them and store alongside your preference dataset. The latter is common for large-scale training — computing reference log probs once saves significant compute.

Interview Question 3: “When Does RLHF Still Win?”

DPO is simpler and often performs comparably. So why would anyone still use RLHF?

Online vs. offline learning. DPO is purely offline — it learns from a fixed preference dataset. RLHF with PPO can generate new samples during training, explore the response space, and potentially find better solutions than exist in your training data.

This matters when your preference data is limited or doesn’t cover the distribution you care about. If you have 10K preference pairs but want to deploy on a diverse set of prompts, RLHF’s ability to explore might help. DPO can only learn from what’s in the dataset.

There’s also the question of reward model interpretability. With RLHF, you can inspect your reward model, probe what it learned, and debug why certain responses score high. With DPO, the reward is implicit — you can extract it, but it’s harder to debug systematically.

My best guess is that for most practical applications, DPO wins on simplicity without sacrificing much quality. But if you’re doing cutting-edge alignment research where exploration matters, RLHF still has advantages.

A sleek, custom-modified white sports car parked on an urban street. Perfect for enthusiasts.
Photo by DUY DAT DANG on Pexels

Interview Question 4: “How Do You Handle Length Bias?”

Both methods suffer from it, but in different ways.

In RLHF, reward models notoriously favor longer responses. The model learns to pad outputs with hedging language, disclaimers, and repetition. You end up with responses that score high on the reward model but are frustrating to read.

DPO inherits this if your preference data has length bias. But it also has a more subtle issue: the log-probability ratio naturally correlates with length. Longer sequences have more tokens, and each token contributes to the log probability. A common fix is length normalization:

def length_normalized_dpo_loss(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    ref_chosen_logps: torch.Tensor,
    ref_rejected_logps: torch.Tensor,
    chosen_lengths: torch.Tensor,
    rejected_lengths: torch.Tensor,
    beta: float = 0.1,
) -> torch.Tensor:
    """
    Length-normalized DPO. Dividing by sequence length prevents
    the model from gaming the loss by generating longer outputs.
    """
    # Average log prob per token
    chosen_ratio = (policy_chosen_logps - ref_chosen_logps) / chosen_lengths
    rejected_ratio = (policy_rejected_logps - ref_rejected_logps) / rejected_lengths

    logits = beta * (chosen_ratio - rejected_ratio)
    loss = -F.logsigmoid(logits).mean()

    return loss

This isn’t in the original DPO paper but shows up in most production implementations. I’ve seen teams skip this and then wonder why their model suddenly becomes verbose after training.

Interview Question 5: “Explain the β\beta Parameter”

In both RLHF and DPO, β\beta controls how much the model can deviate from the reference policy.

  • High β\beta (e.g., 0.5): strong regularization, model stays close to reference, conservative changes
  • Low β\beta (e.g., 0.01): weak regularization, model can drift far, more aggressive optimization

The tricky part: the “right” value depends on your data quality and how different you want the final model to be. I’ve seen β=0.1\beta = 0.1 work well for general instruction tuning, but safety fine-tuning sometimes needs higher values to prevent catastrophic forgetting of safety behaviors.

Here’s a quick experiment showing how β\beta affects training dynamics:

import numpy as np

def simulate_dpo_gradient(beta_values, chosen_ratio, rejected_ratio):
    """
    Show how beta affects gradient magnitude.
    Larger beta means smaller effective gradients when ratios are similar.
    """
    results = []
    for beta in beta_values:
        logit = beta * (chosen_ratio - rejected_ratio)
        # Gradient of logsigmoid is (1 - sigmoid(x))
        grad_scale = 1 - 1 / (1 + np.exp(-logit))
        results.append((beta, logit, grad_scale))
    return results

# Typical early training: chosen and rejected have similar log ratios
for beta, logit, grad in simulate_dpo_gradient([0.01, 0.1, 0.5], 0.5, 0.3):
    print(f"beta={beta}: logit={logit:.3f}, grad_scale={grad:.3f}")

Output:

beta=0.01: logit=0.002, grad_scale=0.500
beta=0.1: logit=0.020, grad_scale=0.495
beta=0.5: logit=0.100, grad_scale=0.475

With low β\beta, gradients stay near 0.5 (maximum) even when the model is close to optimal. High β\beta causes gradients to drop faster as the model improves. This affects convergence speed and final solution quality in ways that aren’t always intuitive.

The Memory Comparison Interviewers Love

Peak GPU memory during training (7B parameter model):

Method Models in Memory Approximate VRAM
RLHF (PPO) Policy + Reference + Reward + Value ~40GB
DPO Policy + Reference ~20GB
DPO (precomputed refs) Policy only ~12GB

DPO’s memory advantage is real but often overstated. If you precompute reference log probabilities, you only need the policy model during training. But you still need to run inference on the reference model once per training example, so total compute isn’t always lower — it’s just shifted.

IPO and KTO: The DPO Variants You Should Know

Interviewers sometimes ask about DPO alternatives. Two worth knowing:

IPO (Identity Preference Optimization) from Azar et al. (2023) addresses a theoretical issue with DPO: as training progresses and the model becomes confident, DPO gradients vanish. IPO uses a different loss that maintains gradients:

LIPO=(log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)−12β)2\mathcal{L}_{IPO} = \left(\log \frac{\pi_\theta(y_w|x)}{\pi_{ref}(y_w|x)} – \log \frac{\pi_\theta(y_l|x)}{\pi_{ref}(y_l|x)} – \frac{1}{2\beta}\right)^2

KTO (Kahneman-Tversky Optimization) works with unpaired preference data — you just need “good” and “bad” examples, not explicit head-to-head comparisons. This is huge for practical data collection, since getting comparison pairs is expensive.

I haven’t tested KTO at scale myself, so take this with a grain of salt: early reports suggest it works nearly as well as DPO with much simpler data requirements.

Real Training Loop: Where Things Actually Break

Here’s a minimal but realistic DPO training loop. The comments highlight where I’ve seen things go wrong:

from transformers import AutoModelForCausalLM, AutoTokenizer
from torch.utils.data import DataLoader
import torch

def train_dpo(
    model_name: str,
    preference_data: list,  # [{"prompt": ..., "chosen": ..., "rejected": ...}]
    beta: float = 0.1,
    epochs: int = 1,
    lr: float = 1e-6,  # DPO often needs very low LR
):
    device = "cuda" if torch.cuda.is_available() else "cpu"

    # Load policy and reference (same init, but reference is frozen)
    policy = AutoModelForCausalLM.from_pretrained(model_name).to(device)
    ref_model = AutoModelForCausalLM.from_pretrained(model_name).to(device)
    ref_model.eval()
    for param in ref_model.parameters():
        param.requires_grad = False

    tokenizer = AutoTokenizer.from_pretrained(model_name)
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token  # Common gotcha

    optimizer = torch.optim.AdamW(policy.parameters(), lr=lr)

    for epoch in range(epochs):
        for batch in preference_data:
            # Tokenize chosen and rejected with prompt
            chosen_ids = tokenizer(
                batch["prompt"] + batch["chosen"],
                return_tensors="pt",
                truncation=True,
                max_length=2048
            ).to(device)

            rejected_ids = tokenizer(
                batch["prompt"] + batch["rejected"],
                return_tensors="pt",
                truncation=True,
                max_length=2048
            ).to(device)

            # Get log probs (this is the tedious part)
            with torch.no_grad():
                ref_chosen_logps = get_sequence_logprob(ref_model, chosen_ids)
                ref_rejected_logps = get_sequence_logprob(ref_model, rejected_ids)

            policy_chosen_logps = get_sequence_logprob(policy, chosen_ids)
            policy_rejected_logps = get_sequence_logprob(policy, rejected_ids)

            loss, _, _ = dpo_loss(
                policy_chosen_logps,
                policy_rejected_logps,
                ref_chosen_logps,
                ref_rejected_logps,
                beta=beta
            )

            optimizer.zero_grad()
            loss.backward()

            # Gradient clipping is essential — I've seen NaNs without it
            torch.nn.utils.clip_grad_norm_(policy.parameters(), 1.0)

            optimizer.step()

    return policy


def get_sequence_logprob(model, inputs):
    """
    Sum of log probs for all tokens in the response (excluding prompt).
    This implementation is simplified — production code handles prompt masking.
    """
    outputs = model(**inputs)
    logits = outputs.logits[:, :-1, :]  # Shift for next-token prediction
    labels = inputs["input_ids"][:, 1:]

    log_probs = torch.nn.functional.log_softmax(logits, dim=-1)
    token_log_probs = log_probs.gather(-1, labels.unsqueeze(-1)).squeeze(-1)

    # Mask padding tokens
    mask = (labels != model.config.pad_token_id).float()
    return (token_log_probs * mask).sum(dim=-1)

The get_sequence_logprob function is where most bugs hide. You need to properly handle prompt masking (only score response tokens), padding, and the off-by-one indexing between logits and labels. TRL library handles this correctly if you’d rather not debug it yourself.

What Most Tutorials Get Wrong

The biggest misconception I see: “DPO is just simpler RLHF.” It’s not. They solve different optimization problems that happen to converge to the same solution under specific assumptions.

RLHF maximizes expected reward with a KL constraint. DPO maximizes the likelihood that the model’s implicit preferences match human preferences. These are mathematically equivalent when the reward follows the Bradley-Terry model and you use the optimal policy form — but the training dynamics differ.

And sometimes DPO’s assumptions break down. If your preference data violates the Bradley-Terry assumption (e.g., preferences are intransitive), DPO can learn weird things. RLHF with a flexible reward model might handle this better, though I’m not entirely sure — I haven’t tested this systematically.

Debugging sessions at 2am are when you really appreciate the difference between these approaches. If you’re going to be up late staring at loss curves, at least have some dark chocolate espresso beans handy.

FAQ

Q: Can I use DPO without a reference model?

No. The reference model is mathematically required to compute the log-ratio that defines DPO’s implicit reward. Some implementations precompute reference log probabilities to avoid loading the model during training, but you still need it at some point. Removing the reference entirely would change the method fundamentally.

Q: Is DPO always better than RLHF for fine-tuning?

Not always. DPO is simpler and uses less memory, making it practical for smaller teams. But RLHF can explore beyond your training data through online sampling, which matters when your preference dataset is limited or doesn’t cover your deployment distribution. For most production use cases with decent preference data, DPO is the pragmatic choice.

Q: What β\beta value should I start with for DPO?

Start with β=0.1\beta = 0.1 for general instruction tuning. If your model drifts too far from the reference (incoherent outputs, forgotten capabilities), increase to 0.2-0.5. If it barely changes from the reference, try 0.05. Safety fine-tuning often needs higher values (0.3-0.5) to preserve existing safety behaviors.

Pick Your Method

Use DPO when you have good preference data and want straightforward training. It’s easier to debug, uses half the memory, and in most benchmarks performs comparably to RLHF. The TRL library from Hugging Face makes implementation trivial.

Use RLHF when you need online exploration, your preference data is sparse, or you want an interpretable reward model you can probe and debug. The complexity tax is real, but sometimes worth paying.

What I’m still curious about: how well do these methods scale beyond current model sizes? The implicit reward in DPO assumes a specific functional form — does that assumption hold as we push toward frontier models with emergent capabilities? The DPO paper showed strong results up to 7B parameters, but systematic comparisons at 70B+ are still scarce. That gap in the literature bothers me.

Did you find this helpful?

Your support keeps this blog running and ad-free content coming.

☕ Buy me a coffee
TODAY 373 | TOTAL 126,581