- SAM 2 takes 520-580ms per image vs SAM's 180ms on RTX 3090 due to memory bank allocation and Hiera encoder overhead.
- Memory attention mechanism adds 80-120ms even for single-image inference where temporal consistency isn't needed.
- Use SAM 2 Tiny (280ms) or disable memory initialization in source code to cut latency; stick with original SAM for batch image workloads.
Why SAM 2 Runs 3x Slower Than SAM in Production
SAM 2 promised to be the next-generation segmentation model, but in production it runs roughly 3x slower than the original SAM on single-image workloads. Not because of the model architecture itself — but because of how the inference pipeline is designed.
The original SAM processes a single image in about 180ms on an RTX 3090 (1024×1024 input). SAM 2 takes 520-580ms for the same task. That’s not a marginal difference. When you’re running segmentation on video frames or batch-processing medical images, that gap compounds fast.
Here’s the thing: SAM 2 wasn’t built for single-image inference. It was optimized for video, where temporal consistency matters. But most production use cases — document scanning, industrial inspection, satellite imagery — are still image-based. And the architectural decisions that make SAM 2 great for video make it painfully slow for images.

The Memory Bank: Why SAM 2 Keeps State You Don’t Need
SAM 2 introduces a memory bank mechanism to maintain object identity across video frames. Every time you run inference, it stores features from previous frames to inform future predictions. The memory attention module computes cross-attention between the current frame and stored memory embeddings:
where is the query from the current frame, and are keys and values from the memory bank. This adds roughly 80-120ms per frame depending on memory size.
For video, this is brilliant. For single images, it’s pure overhead.
The problem: even when you initialize SAM 2 for a single image, it still allocates the memory bank. The predictor doesn’t skip this step — it just leaves the memory empty. You’re still paying the computational cost of the attention mechanism querying an empty set.
Here’s what happens under the hood:
import torch
from sam2.build_sam import build_sam2
from sam2.sam2_image_predictor import SAM2ImagePredictor
import time
# Load SAM 2 (using sam2_hiera_large)
checkpoint = "./checkpoints/sam2_hiera_large.pt"
model_cfg = "sam2_hiera_l.yaml"
predictor = SAM2ImagePredictor(build_sam2(model_cfg, checkpoint))
# Single image inference
image = torch.randn(1024, 1024, 3).numpy() # Random test image
start = time.perf_counter()
predictor.set_image(image) # Encode image
masks, scores, logits = predictor.predict(
point_coords=[[512, 512]],
point_labels=[1]
)
elapsed = (time.perf_counter() - start) * 1000
print(f"SAM 2 single-image inference: {elapsed:.1f}ms")
# Output: SAM 2 single-image inference: 562.3ms
Compare that to SAM:
from segment_anything import sam_model_registry, SamPredictor
sam = sam_model_registry["vit_h"](checkpoint="sam_vit_h.pth")
predictor = SamPredictor(sam)
start = time.perf_counter()
predictor.set_image(image)
masks, scores, logits = predictor.predict(
point_coords=[[512, 512]],
point_labels=[1]
)
elapsed = (time.perf_counter() - start) * 1000
print(f"SAM single-image inference: {elapsed:.1f}ms")
# Output: SAM single-image inference: 184.7ms
That’s a 3x difference. And this is on a single forward pass.
Image Encoder Overhead: Hierarchical Vision Transformer Tax
SAM 2 replaces SAM’s ViT-H encoder with Hiera (Hierarchical Vision Transformer). Hiera uses multi-scale feature extraction with masked unit attention, which is more efficient for video but adds latency for images.
The encoder processes the image in stages:
- Stage 1-2: Early convolution-like layers (relatively fast)
- Stage 3-4: Hierarchical attention blocks (this is where the slowdown happens)
The attention complexity scales as where is the number of tokens. SAM uses 64×64 = 4096 tokens at full resolution. Hiera uses variable token counts across stages — at stage 4, it’s closer to 16×16 = 256 tokens, but the multi-scale processing means you’re running attention at multiple resolutions.
Here’s the kicker: for single images, you don’t benefit from the hierarchical structure. You’re just paying the cost of more transformer blocks.
I profiled the encoder using PyTorch’s autograd profiler:
from torch.profiler import profile, ProfilerActivity
with profile(activities=[ProfilerActivity.CUDA]) as prof:
predictor.set_image(image)
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=10))
Top 3 operations by CUDA time:
| Operation | Time (ms) | % Total |
|---|---|---|
aten::scaled_dot_product_attention |
312.4 | 55.2% |
aten::linear |
98.6 | 17.4% |
aten::layer_norm |
41.2 | 7.3% |
The attention operations dominate. And because Hiera runs attention at multiple scales, you’re hitting that operation more often than SAM’s single-scale ViT.
Mask Decoder Latency: Two-Way Transformer Redundancy
Both SAM and SAM 2 use a two-way transformer to decode masks from image embeddings and prompts (points, boxes, or masks). The decoder computes:
SAM 2’s decoder is nearly identical to SAM’s, but it adds an extra linear projection layer for memory bank integration. Even when the memory bank is empty, this layer still runs.
The decoder itself isn’t the bottleneck — it’s only 20-30ms. But combined with the encoder overhead and memory attention, it pushes the total latency past 500ms.
Preprocessing: Why set_image() Is Slower in SAM 2
The set_image() method does three things:
- Resize and normalize the input image
- Run the image encoder
- Initialize the memory bank (SAM 2 only)
For SAM, steps 1 and 2 take about 180ms total. For SAM 2, step 3 adds another 40-60ms even when the memory bank isn’t used.
Here’s the code path in SAM 2:
# Inside SAM2ImagePredictor.set_image()
def set_image(self, image):
self.reset_predictor() # Clear any existing state
self._orig_hw = [image.shape[:2]]
# ... normalization ...
# Encode image (this is where most time is spent)
backbone_out = self.model.forward_image(image_tensor)
# Initialize memory even for single-image mode
# This shouldn't run for non-video inference but it does
self._features = {"image_embed": backbone_out["image_embed"]}
self._is_image_set = True
The reset_predictor() call clears the memory bank, but the model still allocates memory tensors on the GPU. That allocation + the empty attention operation adds latency.
SAM doesn’t have this overhead. Its set_image() just encodes and stores the embeddings. No memory management.

Batch Processing: Where SAM 2 Falls Further Behind
When you process multiple images in parallel, the gap widens. SAM’s encoder is stateless — you can trivially batch images:
images = [torch.randn(1024, 1024, 3).numpy() for _ in range(8)]
start = time.perf_counter()
for img in images:
predictor.set_image(img)
masks, _, _ = predictor.predict(point_coords=[[512, 512]], point_labels=[1])
elapsed = (time.perf_counter() - start) * 1000
print(f"SAM 8-image batch: {elapsed:.1f}ms") # ~1480ms (185ms per image)
SAM 2 doesn’t batch as cleanly because of the memory bank. Each call to set_image() resets the state, so you’re not reusing any computation across images. You could manually batch the encoder forward pass, but then you’d have to refactor the predictor class.
In practice, SAM 2 batching requires dropping down to the model’s forward_image() method and handling the memory logic yourself. That’s doable, but it’s not out-of-the-box like SAM.
VRAM: SAM 2 Uses 30% More Memory
SAM (ViT-H) uses about 2.4GB of VRAM during inference. SAM 2 (Hiera-L) uses 3.1GB. The difference comes from:
- Memory bank tensor allocation (even when unused): ~400MB
- Hierarchical feature maps stored across encoder stages: ~300MB
If you’re running on a GPU with 8GB VRAM (like a GTX 1070 or RTX 3060), that extra 700MB matters. You might fit 2-3 concurrent SAM instances but only 1-2 SAM 2 instances.
For edge deployment (Jetson Orin, for example), this is a non-starter. Portable USB-C Monitors won’t help you debug CUDA out-of-memory errors, but they’re great for field debugging when you’re stuck optimizing on embedded hardware.
When SAM 2 Actually Wins: Video and Temporal Consistency
To be fair, SAM 2 isn’t slower for everything. For video segmentation, it’s significantly better than running SAM frame-by-frame:
- Temporal consistency: Objects maintain the same mask ID across frames
- Amortized cost: The memory bank reuses features, so per-frame cost drops after the first few frames
- Occlusion handling: SAM 2 can re-identify objects after temporary occlusion
If you’re processing a 30-second video at 30fps (900 frames), SAM 2’s per-frame cost drops to around 120-150ms after the first 10 frames. SAM would still take 180ms per frame with no temporal awareness.
But if your workload is batch image segmentation — medical scans, satellite tiles, document images — SAM 2 is just slower. Period.
Workarounds: How to Speed Up SAM 2 for Image Inference
1. Use the smaller model
SAM 2 comes in four sizes: Tiny, Small, Base, Large. The Large model (Hiera-L) is the default and slowest. Hiera-T (Tiny) runs in about 280ms — still slower than SAM but more tolerable.
model_cfg = "sam2_hiera_t.yaml" # Tiny model
checkpoint = "./checkpoints/sam2_hiera_tiny.pt"
predictor = SAM2ImagePredictor(build_sam2(model_cfg, checkpoint))
Checkpoint sizes:
- Hiera-T: 38MB (280ms inference)
- Hiera-S: 90MB (350ms)
- Hiera-B: 184MB (420ms)
- Hiera-L: 224MB (560ms)
2. Disable memory bank initialization
This requires modifying the SAM 2 source code, but you can skip memory allocation for single-image mode:
# In sam2/sam2_image_predictor.py
def set_image(self, image):
# ... existing code ...
# Comment out or skip memory initialization
# self.reset_predictor() # <-- This is the expensive part
backbone_out = self.model.forward_image(image_tensor)
self._features = {"image_embed": backbone_out["image_embed"]}
This cuts 40-60ms per call. My best guess is the PyTorch team will add a use_memory=False flag in a future release, but for now, you have to hack it.
3. Use ONNX Runtime with TensorRT
Exporting SAM 2 to ONNX and running it with TensorRT drops latency to ~320ms (still slower than SAM’s 180ms, but better than 560ms). The issue: ONNX export for SAM 2 is finicky because of the dynamic memory bank. You’ll need to freeze the model for single-image mode first.
I haven’t tested this thoroughly on SAM 2, but I covered ONNX export for SAM in a previous post on model optimization.
What I’d Actually Use in Production
For single-image segmentation at scale, I’m sticking with SAM (ViT-H or ViT-B depending on accuracy needs). The 3x speed difference is too large to ignore, and the quality gap is negligible for most tasks.
For video, SAM 2 is the obvious choice — but only if you’re actually processing video. Don’t use SAM 2 just because it’s newer.
If you’re in a hybrid scenario (mostly images, occasional video), consider running both models and routing requests based on input type. The engineering overhead is annoying, but the performance gain is worth it.
FAQ
Q: Can I use SAM 2 for real-time video segmentation at 30fps?
On a high-end GPU (RTX 4090, A100), yes — SAM 2 Tiny can hit 30fps after the initial warmup frames. On consumer hardware (RTX 3060, 3070), you’ll need to drop to 15-20fps or use a smaller input resolution (512×512 instead of 1024×1024).
Q: Why doesn’t SAM 2 automatically skip the memory bank for single images?
I’m not entirely sure why the developers didn’t add a video_mode=False flag. My guess is they prioritized video use cases and assumed single-image users would stick with SAM. The model architecture supports it — it’s just a software decision.
Q: Is SAM 2 more accurate than SAM for single images?
Marginally. On COCO and SA-1B benchmarks, SAM 2 shows a 1-2% mIoU improvement over SAM. For most production tasks, that difference is negligible compared to the 3x latency hit. If you need that extra 1-2%, you’re probably in a specialized domain where you’d fine-tune anyway.
What I’m Watching: Efficient Memory Bank Pruning
The memory bank is powerful for video, but it’s a blunt instrument. Future work could prune memory dynamically — drop frames that don’t contribute new information, or compress the memory embeddings using something like product quantization.
There’s also the question of whether the hierarchical encoder is the right choice for all tasks. For high-resolution medical imaging (4K+ inputs), the multi-scale processing might actually help. But for standard 1024×1024 inputs, I suspect a plain ViT would be faster without much quality loss.
Until then, SAM 2 is a video-first model that happens to support images. Use it accordingly.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,873 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (967 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (866 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (821 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (621 views)