BERT vs RoBERTa vs DistilBERT: GLUE Scores Decoded

Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links in this post are affiliate links — they cost you nothing extra.
⚡ Key Takeaways
  • RoBERTa improves BERT by 4+ GLUE points using identical architecture but better training: no NSP, dynamic masking, 10x more data, larger batches.
  • DistilBERT retains 97% of BERT's performance at 60% inference speed and 40% smaller size through knowledge distillation.
  • For new projects: use RoBERTa-base for accuracy, DistilBERT for latency-constrained production, skip BERT-base entirely.

The 88.9% That Started Everything

RoBERTa hit 88.5% on the GLUE benchmark in July 2019. BERT, released just 8 months earlier, scored 80.5% with BERT-base and about 84.5% with BERT-large. That 4-point jump from “large” to RoBERTa came from exactly zero architectural changes.

Wait, what?

Yes, RoBERTa (Liu et al., 2019) is literally BERT with better training. Same transformer encoder, same masked language modeling objective (mostly), same attention mechanism. The gains came from training longer, on more data, with bigger batches, and without the next sentence prediction (NSP) task that BERT insisted was important.

This post walks through what actually changed between these three models, why the GLUE numbers moved the way they did, and which one I’d actually deploy today.

Cup of coffee and a city map on a wooden table by the window in Copenhagen's cozy cafe.
Photo by Berna Deniz on Pexels

What Is GLUE and Why Should You Care?

The General Language Understanding Evaluation benchmark bundles 9 NLP tasks into a single score. It includes sentiment analysis (SST-2), sentence similarity (STS-B, MRPC, QQP), natural language inference (MNLI, RTE, QNLI), and linguistic acceptability (CoLA). A model gets scored on each task, and the GLUE score is roughly the average.

Before BERT, the state-of-the-art hovered around 69%. BERT-large pushed it to 82.1% on the public leaderboard. Then RoBERTa hit 88.5%. These jumps mattered because they showed pre-trained language models could crush benchmarks that previously required task-specific architectures.

But here’s what nobody tells beginners: GLUE saturated fast. By late 2019, models were hitting 90%+ and the benchmark stopped being useful for distinguishing top performers. SuperGLUE replaced it. Still, GLUE remains the canonical comparison point for understanding BERT-family models.

Enjoying this article? Get more like it delivered to your inbox. Subscribe to the newsletter

BERT: The Original Recipe

Devlin et al. (2019) introduced BERT with two training objectives:

  1. Masked Language Modeling (MLM): Randomly mask 15% of tokens, predict them. The loss is simply:

LMLM=−∑i∈maskedlog⁡P(xi∣x\i)L_{MLM} = -\sum_{i \in \text{masked}} \log P(x_i | x_{\backslash i})

  1. Next Sentence Prediction (NSP): Given two sentences A and B, predict whether B actually follows A in the corpus. Binary cross-entropy:

LNSP=−[ylog⁡(p)+(1−y)log⁡(1−p)]L_{NSP} = -[y \log(p) + (1-y)\log(1-p)]

BERT-base: 12 layers, 768 hidden, 12 heads, 110M parameters. BERT-large: 24 layers, 1024 hidden, 16 heads, 340M parameters. Training used BooksCorpus (800M words) plus English Wikipedia (2.5B words) for about 1M steps with batch size 256.

The GLUE results from the original paper:

Model MNLI QQP QNLI SST-2 CoLA STS-B MRPC RTE Avg
BERT-base 84.6 71.2 90.5 93.5 52.1 85.8 88.9 66.4 79.1
BERT-large 86.7 72.1 92.7 94.9 60.5 86.5 89.3 70.1 82.1

CoLA stands out as the hardest—a 60.5% on linguistic acceptability meant the model still struggled with subtle grammatical judgments.

RoBERTa: Same Architecture, Better Training

Liu et al. at Facebook AI didn’t change the transformer. They changed everything else.

Here’s the recipe:

  1. Remove NSP. Turns out it didn’t help. In fact, it might have hurt. Training on individual sentences without the NSP objective improved downstream performance.

  2. Dynamic masking. BERT generated masks once during preprocessing and reused them. RoBERTa regenerates masks each time a sequence is fed to the model. More diversity during training.

  3. More data. BERT used ~16GB of text. RoBERTa used ~160GB: CC-News, OpenWebText, Stories corpus, plus the original BERT data.

  4. Longer training. BERT trained for 1M steps. RoBERTa trained for 500K steps but with batch size 8K instead of 256. That’s 16x more samples per step.

  5. Larger batches. They found that larger batch sizes (2K, 4K, 8K) with appropriately scaled learning rates improved performance.

The total compute was roughly 10x BERT’s. But the architecture? Identical to BERT-large.

RoBERTa-large GLUE scores:

Task RoBERTa BERT-large Δ
MNLI 90.2 86.7 +3.5
QQP 72.2 72.1 +0.1
QNLI 94.7 92.7 +2.0
SST-2 96.4 94.9 +1.5
CoLA 68.0 60.5 +7.5
STS-B 92.4 86.5 +5.9
MRPC 90.9 89.3 +1.6
RTE 86.6 70.1 +16.5

That RTE jump is wild. From 70% to 86% on natural language inference just from training better? I’m not entirely sure why RTE specifically benefited so much—my best guess is the additional data covered more inference patterns, and the longer training allowed the model to generalize them.

DistilBERT: Half the Size, 97% of the Performance

Sanh et al. (2019) at Hugging Face asked a different question: what if we compress BERT instead of scaling it?

DistilBERT uses knowledge distillation. You take a trained “teacher” model (BERT-base) and train a smaller “student” model to mimic its outputs. The student learns not just the hard labels but the teacher’s probability distribution over the vocabulary.

The distillation loss combines three terms:

Ldistil=α⋅LCE(y,ps)+β⋅LKL(pt,ps)+γ⋅Lcos(ht,hs)L_{distil} = \alpha \cdot L_{CE}(y, p_s) + \beta \cdot L_{KL}(p_t, p_s) + \gamma \cdot L_{cos}(h_t, h_s)

Where LCEL_{CE} is cross-entropy with true labels, LKLL_{KL} is KL-divergence between teacher and student softmax outputs (with temperature TT), and LcosL_{cos} is cosine embedding loss between hidden states.

The temperature-softened probability comes from:

pi=exp⁡(zi/T)∑jexp⁡(zj/T)p_i = \frac{\exp(z_i/T)}{\sum_j \exp(z_j/T)}

Higher temperature makes the distribution softer, revealing more information about the teacher’s uncertainty.

DistilBERT architecture: 6 layers (half of BERT-base’s 12), same hidden size (768), 66M parameters (40% smaller). Inference is about 60% faster.

GLUE comparison:

Model Params Speed Avg GLUE
BERT-base 110M 1.0x 79.1
DistilBERT 66M 1.6x 76.8
RoBERTa-base 125M 1.0x 86.4

DistilBERT retains about 97% of BERT-base’s performance at 60% the size. That’s the headline number. But look at the absolute gap: 76.8 vs 79.1 is 2.3 points. On CoLA specifically, DistilBERT drops from 52.1 to 49.0—a bigger relative decline.

The Training Details That Trip People Up

If you’re fine-tuning these models, here’s what the papers don’t emphasize enough.

Learning rate sensitivity. BERT recommends fine-tuning with learning rates in {2e-5, 3e-5, 5e-5}. RoBERTa used 1e-5 for some tasks. Using 5e-5 on RoBERTa can destabilize training on small datasets like RTE (2.5K examples). I’d recommend starting at 1e-5 for RoBERTa and 2e-5 for BERT/DistilBERT.

Warmup matters. All three papers use linear warmup. RoBERTa specifically used 6% of training steps for warmup. Skip warmup and you’ll see gradient explosions in the first 100 steps.

Batch size tradeoffs. The papers used batch sizes of 16 or 32 for fine-tuning. But if your GPU can’t fit 32, gradient accumulation works fine. Effective batch size of 32 via 4 accumulation steps ≈ actual batch size 32 in most cases. (Though for very small datasets, the stochasticity of smaller batches might actually help.)

Epochs. BERT paper recommends 2-4 epochs. But here’s the thing: on tiny datasets like RTE, 10 epochs sometimes works better. I’ve seen people religiously stick to 3 epochs and underfit.

Here’s a minimal fine-tuning setup that actually works:

from transformers import AutoModelForSequenceClassification, AutoTokenizer
from transformers import TrainingArguments, Trainer
import torch

model_name = "roberta-base"  # or "bert-base-uncased", "distilbert-base-uncased"
model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=2)
tokenizer = AutoTokenizer.from_pretrained(model_name)

# check your versions - this works on transformers 4.35+
print(f"transformers version: {transformers.__version__}")

training_args = TrainingArguments(
    output_dir="./results",
    learning_rate=1e-5,  # lower for RoBERTa
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,  # effective batch = 32
    num_train_epochs=4,
    warmup_ratio=0.06,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    metric_for_best_model="eval_accuracy",
)

# the usual Trainer setup...
# watch for this warning on newer transformers:
# "Some weights of RobertaForSequenceClassification were not initialized"
# That's expected - the classification head is randomly initialized

The tokenizer gotcha. BERT uses [CLS] and [SEP]. RoBERTa uses <s> and </s>. DistilBERT uses BERT’s tokens. If you’re manually tokenizing (why?), getting these wrong silently degrades performance. Just use AutoTokenizer and let it handle it.

A green metal wall featuring an emergency exit sign in Dutch. Perfect for safety-themed graphics.
Photo by Jan van der Wolf on Pexels

Which Ablation Surprised Me Most?

The RoBERTa paper has a fascinating ablation on NSP. They tested four configurations:

  1. Segment-pair + NSP (original BERT)
  2. Sentence-pair + NSP (pairs of actual sentences)
  3. Full sentences without NSP (pack documents, no sentence breaks)
  4. Document sentences without NSP (don’t cross document boundaries)

Option 4 performed best. But here’s the surprise: option 1 (BERT’s approach) was the worst among them. The NSP task combined with segment pairs actually hurt the model.

Why? My interpretation: NSP is a trivially easy task (97%+ accuracy quickly), so it doesn’t provide useful gradients. Meanwhile, the segment pair setup fragments input weirdly. Documents have flow; artificially pairing random segments breaks that.

This matters because it shows how much performance we left on the table with BERT’s original design. Not because BERT was bad—it was revolutionary—but because the authors couldn’t run every experiment.

Practical Performance: What the Numbers Actually Mean

Let me be direct about deployment.

DistilBERT is the right choice for 90% of production NLP. The 2-point GLUE drop translates to maybe 1-2% accuracy drop on most real tasks. But you get 1.6x inference speed and 40% smaller memory footprint. On my M1 MacBook (using MPS backend), DistilBERT processes a batch of 32 sentences in ~45ms. BERT-base takes ~72ms. RoBERTa-base is similar to BERT.

RoBERTa-base makes sense when you have abundant fine-tuning data and accuracy matters more than latency. If you’re building a content moderation system processing millions of posts daily, that extra accuracy point might catch thousands of edge cases.

BERT-base is mostly for legacy compatibility now. If you’re starting fresh, RoBERTa-base gives you better performance with identical inference cost.

What about the “large” variants? On a CPU, BERT-large inference is roughly 3x slower than BERT-base. On an A100 with batching, the gap shrinks to maybe 1.5x. If you need maximum accuracy and have GPU budget, RoBERTa-large is the pick. But DistilBERT often beats BERT-large on latency-normalized accuracy.

Debugging at 2am while comparing model variants? Dark Chocolate Espresso Beans help more than you’d think.

What the Authors Acknowledged (and What They Didn’t)

The RoBERTa paper is refreshingly honest: “We find that BERT was significantly undertrained.” They explicitly state that with sufficient compute, BERT’s architecture is competitive. The gains are training, not architecture.

DistilBERT’s limitation is more subtle. The paper claims 97% performance retention, but that varies by task. On CoLA, retention drops to 94%. On tasks requiring deep linguistic reasoning, distillation loses more. The authors mention this briefly but don’t deeply analyze why some tasks compress better than others.

What none of them discuss well: domain shift. All GLUE tasks are well-edited English text. How do these models perform on tweets? Medical records? Legal documents? In practice, domain-specific models (BioBERT, LegalBERT) often beat RoBERTa despite lower general GLUE scores. GLUE is a proxy, not a guarantee.

BERT vs RoBERTa: When Training Beats Architecture

A common question: should I use BERT or RoBERTa for a new project?

RoBERTa. Always RoBERTa. Same architecture, better checkpoint. The only exception is if you need a multilingual model—mBERT exists and is well-supported, while mRoBERTa isn’t widely available.

But here’s the nuance: RoBERTa’s strength comes from its pre-training. If you’re going to pre-train from scratch on domain data (say, 10GB of medical text), the RoBERTa training recipe matters more than the RoBERTa checkpoint. You’d use BERT’s architecture with RoBERTa’s hyperparameters: no NSP, dynamic masking, larger batches, longer training.

I covered the evolution of attention mechanisms in LSTM Attention vs Self-Attention: How Bahdanau Evolved, which explains why the transformer encoder became the de facto choice for this family.

DistilBERT vs BERT-large: The Efficiency Question

This comparison is underrated.

Model Params GLUE Inference (CPU)
BERT-base 110M 79.1 1.0x
BERT-large 340M 82.1 3.0x
DistilBERT 66M 76.8 0.6x

BERT-large is 3x slower than BERT-base for a 3-point improvement. DistilBERT is 0.6x the speed of BERT-base for a 2.3-point drop. If your latency budget is fixed, running two DistilBERT models in an ensemble might give you better accuracy than one BERT-large—and faster.

I haven’t tested this rigorously across all GLUE tasks. Take it with a grain of salt. But the math suggests ensemble small models can outperform single large ones within fixed compute budgets.

FAQ

Q: Can I use RoBERTa as a drop-in replacement for BERT?

Yes, with one caveat: RoBERTa doesn’t have the token_type_ids that BERT uses for segment A/B distinctions. Most modern code handles this automatically, but if you’re using older scripts that explicitly set token_type_ids, you’ll get errors or silent failures. Check that your tokenizer returns None for token_type_ids with RoBERTa models.

Q: Is DistilBERT good enough for production?

For most classification tasks, yes. The 97% performance retention holds up on sentiment analysis, intent detection, and similar tasks. Where DistilBERT struggles: anything requiring deep linguistic reasoning (CoLA-style tasks), long-range dependencies, or nuanced inference. If your use case is straightforward classification, DistilBERT is often the best latency/accuracy tradeoff.

Q: Why didn’t RoBERTa change the architecture if training improvements were so effective?

The RoBERTa paper’s explicit goal was isolating the effect of training decisions. They wanted to prove BERT was undertrained, not that transformers needed redesigning. Architectural innovations came later with models like ELECTRA, DeBERTa, and the efficient transformers (Longformer, BigBird). RoBERTa’s contribution was methodological: showing that scaling laws apply to pre-training.

References

  • Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL. https://arxiv.org/abs/1810.04805

  • Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. https://arxiv.org/abs/1907.11692

  • Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. EMC² Workshop at NeurIPS. https://arxiv.org/abs/1910.01108

  • Wang, A., Singh, A., Michael, J., et al. (2019). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. ICLR. https://arxiv.org/abs/1804.07461


For new projects: RoBERTa-base for accuracy, DistilBERT for speed, forget BERT-base exists. The architectural debates have moved on to DeBERTa, efficient attention mechanisms, and decoder-only models. But understanding why RoBERTa beat BERT with zero architecture changes? That’s the lesson that keeps paying off—training decisions compound.

What I’m still curious about: DistilRoBERTa exists but never got as popular as DistilBERT. I’d expect distilling from a better teacher to help, but the empirical adoption pattern suggests otherwise. Maybe the community just standardized on DistilBERT before DistilRoBERTa matured, or maybe there’s a real performance gap I haven’t benchmarked.

Did you find this helpful?

Your support keeps this blog running and ad-free content coming.

☕ Buy me a coffee
TODAY 36 | TOTAL 126,728