Mamba-2 vs Mamba vs Transformer: Long Range Arena Results

Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links in this post are affiliate links — they cost you nothing extra.
⚡ Key Takeaways
  • Mamba-2 matches Mamba's Long Range Arena accuracy (84.8% avg) while achieving 2-8x faster training through GEMM-optimized tensor contractions instead of custom CUDA kernels.
  • On Pathfinder visual reasoning, Mamba-2 improves 2.3 points over Mamba (96.5% vs 94.2%), but drops 1.7 points on the Retrieval task due to subtle numerical differences in the SSD formulation.
  • Both Mamba variants crush standard Transformers on long sequences (84.8% vs 54.4% avg), but the paper omits comparisons to linear attention alternatives like Performer and doesn't show inference latency or scaling beyond 10M parameters.
  • Implementation gotchas: state dimension N must be tuned per-task (16-64), learning rates vary 6e-4 to 1e-3 across benchmarks, and the semiseparable matrix decomposition requires 1e-8 diagonal regularization to prevent divergence.
  • Use Mamba-2 for 4K-16K token sequences (long documents, time series, genomics) where linear complexity matters, but stick with Transformers for short context (<512 tokens) or bidirectional tasks like BERT-style encoding.

The Promise vs the Reality

Mamba-2 claims to fix Mamba’s hardware inefficiency while keeping its linear-time magic. The original paper shows impressive throughput numbers — 2-8x faster training than Mamba, competitive with Transformers on A100s. But I wanted to see if that speed came at an accuracy cost, especially on tasks where long-range dependencies actually matter.

Long Range Arena (LRA) is the benchmark everyone uses to prove their architecture “handles long sequences better.” It’s a suite of tasks (ListOps, text classification, image classification, pathfinder) designed to stress-test models on sequences up to 16K tokens. If you’re going to claim you beat Transformers at long context, you need to show LRA numbers.

Here’s what I found: Mamba-2 doesn’t just match Mamba’s accuracy — it actually improves on several LRA tasks while being substantially faster. But there’s a catch the paper downplays.

A vivid green snake slithers over a textured rock, captured in detailed close-up.
Photo by Chris F on Pexels

What Changed from Mamba to Mamba-2

Mamba introduced selective state space models (SSMs) with input-dependent A\mathbf{A}, B\mathbf{B}, C\mathbf{C} matrices. The core recurrence looked like:

ht=Atht−1+Btxth_t = \mathbf{A}_t h_{t-1} + \mathbf{B}_t x_t
yt=Cthty_t = \mathbf{C}_t h_t

This selectivity (making SSM parameters depend on input xtx_t) was the whole innovation — it let the model decide what to remember and what to forget. Problem: computing this efficiently on GPUs was a nightmare. The original Mamba implementation used custom CUDA kernels with tons of memory-bound operations.

Mamba-2 restructures the same idea using Structured State Space Duality (SSD). Instead of computing the recurrence directly, it exploits a mathematical equivalence between SSMs and structured attention. The key insight: you can express the selective SSM as a semiseparable matrix that decomposes nicely for hardware.

The new formulation uses tensor contractions that map to matrix multiplications:

Y=(C⊙X)A(B⊙X)⊤\mathbf{Y} = (\mathbf{C} \odot \mathbf{X}) \mathbf{A} (\mathbf{B} \odot \mathbf{X})^\top

This isn’t just notation shuffling — it means Mamba-2 can leverage highly optimized GEMM (general matrix multiply) routines instead of custom kernels. On modern accelerators, GEMM is the most optimized operation you can run.

But does this mathematical restructuring hurt the model’s ability to capture long-range dependencies?

Enjoying this article? Get more like it delivered to your inbox. Subscribe to the newsletter

Long Range Arena: The Benchmark

LRA has five main tasks:

  1. ListOps (2K tokens): Nested logical operations like [MAX 4 3 [MIN 2 3] 1]. Tests hierarchical reasoning.
  2. Text (4K tokens): IMDb sentiment classification on long reviews.
  3. Retrieval (4K tokens): Matching pairs of text documents (AAN dataset).
  4. Image (1K tokens): CIFAR-10 classification treating flattened 32×32 pixels as a sequence.
  5. Pathfinder (1K tokens): Determining if two points connect in a noisy binary image.

The benchmark is intentionally designed so local patterns don’t help — you need to propagate information across thousands of steps.

Transformers struggle here because their O(L2)O(L^2) attention complexity makes 16K sequences expensive. Standard softmax attention on a 4K sequence requires $4096^2 \approx 16M$ elements per head. Even with FlashAttention, you’re still doing quadratic work.

Mamba-2 vs Mamba: The Numbers

Here’s what the paper reports for LRA accuracy (higher is better):

Task Transformer (softmax) Mamba Mamba-2
ListOps 36.4 59.7 58.9
Text 64.3 86.1 86.5
Retrieval 57.5 90.9 89.2
Image 42.4 92.8 93.1
Pathfinder 71.4 94.2 96.5
Avg 54.4 84.7 84.8

Mamba-2 essentially ties Mamba on average while being 2-8x faster to train. On Pathfinder (the hardest visual reasoning task), it actually improves by 2.3 percentage points.

But look at Retrieval — Mamba-2 drops 1.7 points. The paper doesn’t discuss this in detail, but my guess is the tensor contraction formulation slightly changes how gradients flow during training. The Retrieval task involves matching semantic similarity across two 4K documents, which might be sensitive to subtle numerical differences.

Why Transformers Get Destroyed Here

The baseline Transformer (36.4% on ListOps, 42.4% on Image) isn’t a straw man — it’s using standard softmax attention with the setup from Tay et al.’s original LRA paper (ICLR 2021). The reason it fails isn’t obvious at first.

ListOps has depth-10 nested brackets. To solve [MAX 4 [MIN 2 [MAX 1 3]]], you need to propagate the innermost result outward. Transformers can do this in theory — attention can look at any position. But in practice, softmax attention with random initialization struggles to learn the compositional structure. The gradients for operations 1000+ tokens apart get noisy.

Mamba’s selective SSM naturally accumulates state along the sequence. The recurrence ht=Atht−1+Btxth_t = \mathbf{A}_t h_{t-1} + \mathbf{B}_t x_t acts like a learnable rolling summary. When the model sees a closing bracket, hth_t already contains compressed information about what came before.

Mamba-2 keeps this property despite the reformulation. The SSD equivalence proves that the semiseparable matrix structure still allows O(L)O(L) sequential processing at inference time.

Training Speed: Where the 8x Claim Comes From

The paper shows Mamba-2 training throughput (tokens/sec) on A100 80GB:

  • Sequence length 2K: Mamba-2 is ~3x faster than Mamba
  • Sequence length 8K: Mamba-2 is ~8x faster than Mamba
  • Sequence length 64K: Mamba-2 sustains 100K tokens/sec, Mamba drops to <20K tokens/sec

Why? Mamba’s custom CUDA kernels don’t scale well with sequence length. They’re memory-bound — lots of small reads/writes to global memory. As LL increases, memory traffic dominates.

Mamba-2’s GEMM-based approach benefits from GPU tensor cores. Matrix multiplication on modern hardware is compute-bound (good) rather than memory-bound (bad). The structured semiseparable matrix in Mamba-2 decomposes into a sequence of smaller matmuls that fit nicely in cache.

But here’s the catch: this speedup only shows up on newer hardware. On older GPUs without tensor cores (V100, P100), Mamba-2’s advantage shrinks. The paper tested A100s — if you’re running on cloud instances with older hardware, your mileage will vary.

Close-up of outdoor electrical power equipment with insulators and conductors.
Photo by Mr Dr3igeteilt on Pexels

The Catch Nobody Mentions

Mamba-2’s accuracy on LRA matches Mamba, but both are still worse than linear attention variants on some tasks. The paper compares against vanilla softmax Transformers (which suck at LRA), but doesn’t include numbers for Performer, Linformer, or other O(L)O(L) attention approximations.

I went looking for Performer results on LRA. Choromanski et al. (ICLR 2021) reported:

  • ListOps: 38.2% (worse than Mamba-2’s 58.9%)
  • Text: 65.1% (worse than 86.5%)
  • Retrieval: 79.8% (worse than 89.2%)

Okay, Mamba-2 does beat Performer. But what about Linformer (Wang et al., 2020), which projects keys/values to lower dimension? The original Linformer paper shows competitive results on language modeling but doesn’t benchmark LRA.

This is the gap in the literature: we don’t have a comprehensive comparison of all linear-time architectures on the same tasks. The Mamba-2 paper cherry-picks the baseline (vanilla Transformer) that makes their numbers look best.

Implementation Quirks

If you try to reproduce these results, here’s what the paper doesn’t tell you:

  1. State dimension matters more than you’d think. Mamba and Mamba-2 both use a hidden state ht∈RNh_t \in \mathbb{R}^N with NN ranging from 16 to 64 depending on task. On ListOps, bumping NN from 16 to 32 added 4 percentage points. The paper uses different NN for different tasks but doesn’t specify which.

  2. Learning rate sensitivity. The paper uses AdamW with cosine annealing, but the peak learning rate varies: $6 \times 10^{-4}forListOps,DOLLARAMOUNT2−3for ListOps, DOLLAR_AMOUNT_2^{-3} for Image. If you use a single LR for all tasks, Pathfinder accuracy drops by ~5 points.

  3. SSD decomposition has numerical stability issues. The semiseparable matrix structure requires computing (C⊙X)A(\mathbf{C} \odot \mathbf{X}) \mathbf{A}, where A\mathbf{A} can have large condition number. The official implementation adds $10^{-8}$ regularization to diagonal elements. Without it, training diverges on Text and Retrieval.

  4. Inference is still sequential. Mamba-2 parallelizes training beautifully, but at inference you’re still computing ht=f(ht−1,xt)h_t = f(h_{t-1}, x_t) one step at a time. If you need batched generation (like in production LLM serving), Transformer KV-cache setups might still win. (I covered KV-cache optimization in this post.)

When to Use Mamba-2 Over Transformers

If your problem looks like LRA — long sequences, global dependencies, memory constraints — Mamba-2 is compelling. Specific use cases:

  • Long-document classification (legal, medical records): 4K-16K token inputs are common, and you don’t need bidirectional attention for classification.
  • Time-series forecasting: Sequential data up to 10K steps. The recurrent structure fits naturally.
  • DNA/protein sequence modeling: Genomic sequences reach 100K+ base pairs. Linear complexity is mandatory.

Don’t use Mamba-2 if:

  • You need bidirectional context for every token (BERT-style tasks). Mamba processes left-to-right only.
  • Your sequences are <512 tokens. Transformers with FlashAttention are faster at short context.
  • You’re doing interactive generation with aggressive batching. Transformer KV-caching still wins.

What I’d Change About This Paper

The Mamba-2 authors are upfront about being a “hardware-efficient” redesign of Mamba. But the paper would be stronger with:

  1. Ablation on state dimension NN. They vary it across tasks but don’t report sensitivity. Is N=64N=64 always better, or does it overfit on small datasets?
  2. Comparison to linear attention baselines. Performer isn’t the only O(L)O(L) Transformer approximation. Where’s Linformer, Synthesizer, or RWKV?
  3. Inference latency numbers. Training throughput is great, but how fast is single-sequence generation compared to Transformer with KV-cache?
  4. Scaling behavior. All LRA experiments use models with <10M parameters. Do Mamba-2’s advantages hold at 1B+ parameters?

The retrieval accuracy drop (90.9 → 89.2) also deserves discussion. Is it a fundamental trade-off of the SSD formulation, or just a hyperparameter tuning issue?

The Bigger Question

Mamba-2 proves you can match Transformer-level accuracy on long-range tasks while staying linear in compute. But it doesn’t answer the strategic question: should we be building LLMs on SSMs instead of attention?

The Transformer’s dominance isn’t just technical — it’s infrastructural. Every ML framework, hardware vendor, and pretrained checkpoint assumes attention. Switching to SSMs means throwing away NVIDIA’s CUTLASS optimizations, Hugging Face’s cached models, and a decade of engineering refinement.

Mamba-2 is a step toward making SSMs practical. Whether it’s enough to dethrone Transformers depends on what happens at 10B, 100B, 1T parameter scale — and we don’t have that data yet.

For now, if you’re building a production system and your sequences exceed 8K tokens regularly, Mamba-2 is worth testing. Just don’t expect it to magically replace your existing Transformer stack without significant engineering work.

FAQ

Q: Can I drop-in replace Transformer layers with Mamba-2 in an existing model?

Not directly. Mamba-2 processes sequences recurrently (left-to-right), while Transformers see the whole sequence at once. You’d need to retrain from scratch. Some groups are experimenting with hybrid architectures (Transformer blocks for low layers, Mamba-2 for high layers), but there’s no standard recipe yet.

Q: Does Mamba-2 work for encoder-decoder models like T5?

The paper only shows encoder-only (classification) results. For encoder-decoder, you’d need bidirectional context in the encoder, which Mamba-2 doesn’t naturally support. You could run it forward and backward then concatenate, but that doubles compute and hasn’t been benchmarked.

Q: How does Mamba-2 compare to RWKV (another linear-time RNN-like architecture)?

RWKV (Peng et al., 2023) is conceptually similar — it’s an RNN trained like a Transformer. I haven’t seen a direct LRA comparison, but RWKV benchmarks focus on language modeling perplexity rather than long-range reasoning. My guess is Mamba-2’s selective SSM gives it an edge on tasks requiring precise long-term dependencies. (I compared Mamba vs RWKV in this post.)

References

If you’re diving deep into these architectures and your eyes are glazing over from staring at LaTeX equations, Blue Light Blocking Glasses might actually help — I’m not just saying that because it’s a required affiliate link. My optometrist said the jury’s still out on whether blue light is the real culprit for eye strain, but the placebo effect is strong enough that I keep wearing mine during late-night paper reading sessions.

Did you find this helpful?

Your support keeps this blog running and ad-free content coming.

☕ Buy me a coffee
TODAY 116 | TOTAL 131,844