- YOLO trains fastest (2.1h) but limits architecture flexibility—swapping loss functions or heads requires forking the library
- MMDetection Faster R-CNN achieved 41% better localization ([email protected]: 0.58 vs YOLO's 0.41) in 4.9 hours with modular config-based experimentation
- Detectron2 offers explicit API control but trains 15% slower than MMDetection for identical models—unclear why, likely data pipeline optimization
MMDetection has the highest learning curve I’ve encountered in object detection frameworks. But it’s also the only one I’d trust for a 50-class custom dataset.
Most tutorials will tell you to start with YOLO because it’s “easy.” They’re not wrong—I had a YOLOv8 model running on a custom hardhat detection dataset in under 30 minutes. But three weeks later, when I needed to swap in Cascade R-CNN because single-stage detectors were missing small objects at distance, I was stuck rewriting everything. The YOLO ecosystem is optimized for speed and convenience, not flexibility.
I spent a month training the same construction site safety detection task (5 classes: hardhat, no-hardhat, vest, machinery, person) across all three frameworks. Same dataset, same train/val split, same Mosaic augmentation scheme. The differences weren’t just in mAP—they showed up in training time, debugging pain, and how easily I could swap architectures when the first attempt failed.
Here’s what actually matters when you’re choosing a framework for custom detection work.

YOLO: Fast Prototyping, Limited Architecture Control
The Ultralytics YOLOv8 API is genuinely elegant. Install via pip, write 10 lines, and you’re training:
from ultralytics import YOLO
model = YOLO('yolov8n.pt') # nano model, 3.2M params
results = model.train(
data='hardhat.yaml',
epochs=100,
imgsz=640,
batch=16,
device=0
)
This trains in 2.1 hours on a single RTX 3090 (24GB VRAM) for 2,400 training images at 640×640. The hardhat.yaml config is just paths to your train/val folders and a class list—no need to register datasets or write custom data loaders.
But here’s where YOLO’s simplicity becomes a constraint. The loss function is baked in: where is the Complete IoU regression loss, is binary cross-entropy for classification, and is the distribution focal loss for box refinement. You can’t easily swap in GIoU or Focal Loss without forking the library.
I wanted to experiment with class-balanced loss because my dataset was imbalanced (400 “hardhat” images vs 80 “machinery”). Ultralytics offers cls_pw for class weights, but it’s a single scalar multiplier—not the per-class weighting you’d get with a custom loss. The framework assumes you’re okay with the default design choices, which is fine for 80-90% of projects.
Inference speed is YOLO’s obvious win. YOLOv8n runs at 120 FPS on the 3090 for 640×640 images. Model size is 6.3 MB for the entire checkpoint (compared to 170 MB for a Faster R-CNN ResNet-50). If you’re deploying to edge devices or need real-time video inference, this gap matters.
But let’s talk about a failure mode I hit: YOLO struggled with crowded scenes. At IoU threshold 0.5, mAP was 0.68. At IoU 0.75 (stricter localization), it dropped to 0.41. The single-stage detector was making confident predictions but with loose bounding boxes—common when objects overlap or appear at varying scales.
Detectron2: Facebook’s Research Playground
Detectron2 is what FAIR (Facebook AI Research) uses internally for vision research, and it shows. The API assumes you’re comfortable reading source code and tweaking config files.
First, you register your dataset manually:
from detectron2.data import DatasetCatalog, MetadataCatalog
from detectron2.data.datasets import register_coco_instances
register_coco_instances(
"hardhat_train",
{},
"./annotations/train.json", # COCO format required
"./images/train"
)
Detectron2 only speaks COCO JSON. If your annotations are in YOLO txt format (one file per image with class x_center y_center width height), you’ll need to convert them. I wrote a 60-line script to handle this—not hard, but an extra step.
Training setup is more verbose:
from detectron2 import model_zoo
from detectron2.config import get_cfg
from detectron2.engine import DefaultTrainer
cfg = get_cfg()
cfg.merge_from_file(model_zoo.get_config_file(
"COCO-Detection/faster_rcnn_R_50_FPN_3x.yaml"
))
cfg.DATASETS.TRAIN = ("hardhat_train",)
cfg.DATASETS.TEST = ()
cfg.DATALOADER.NUM_WORKERS = 4
cfg.MODEL.WEIGHTS = model_zoo.get_checkpoint_url(
"COCO-Detection/faster_rcnn_R_50_FPN_3x.yaml"
)
cfg.SOLVER.IMS_PER_BATCH = 4 # effective batch size
cfg.SOLVER.BASE_LR = 0.001
cfg.SOLVER.MAX_ITER = 5000
cfg.MODEL.ROI_HEADS.BATCH_SIZE_PER_IMAGE = 512 # RoI per image
cfg.MODEL.ROI_HEADS.NUM_CLASSES = 5
trainer = DefaultTrainer(cfg)
trainer.resume_or_load(resume=False)
trainer.train()
Training Faster R-CNN R-50 FPN for 5,000 iterations (roughly 35 epochs given my batch size) took 5.7 hours on the same 3090. That’s 2.7x longer than YOLO. Why? Two-stage detectors have an RPN (Region Proposal Network) that generates ~2,000 proposals per image, then an R-CNN head that classifies and refines each proposal. The forward pass is where both components have separate box regression and classification losses: . More computation per image.
But [email protected]:0.95 (COCO metric averaging IoU from 0.5 to 0.95) was 0.52 vs YOLO’s 0.46. At IoU 0.75, Faster R-CNN scored 0.58 vs YOLO’s 0.41—a 41% improvement where precise localization matters.
Detectron2’s real strength is architecture flexibility. I swapped to Cascade R-CNN (which applies three sequential R-CNN stages with increasing IoU thresholds—0.5, 0.6, 0.7—to progressively refine boxes) by changing one line:
cfg.merge_from_file(model_zoo.get_config_file(
"Cascade-RCNN-X-152-32x8d-FPN-IN5k.yaml"
))
Cascade R-CNN bumped [email protected] to 0.63, a further 8.6% gain. Training time jumped to 9.3 hours (ResNeXt-152 backbone), but for a production pipeline where precision matters more than training cost, this flexibility is invaluable.
The catch: Detectron2 development has slowed. Last major release was 2021. The repo is stable but not actively adding new models. If you want Swin Transformer backbones or DINO (the new self-supervised detection approach), you’ll need to integrate them yourself or look elsewhere.
MMDetection: The Swiss Army Knife (with a manual in Chinese)
MMDetection is part of the OpenMMLab ecosystem—a modular framework supporting 50+ detection architectures, from RetinaNet to DETR to modern anchor-free detectors like FCOS and ATSS.
The learning curve is steep. Configuration is done via Python config files that inherit from base configs:
# configs/custom/faster_rcnn_r50_hardhat.py
_base_ = '../faster_rcnn/faster_rcnn_r50_fpn_1x_coco.py'
model = dict(
roi_head=dict(
bbox_head=dict(num_classes=5)
)
)
dataset_type = 'CocoDataset'
data_root = './data/hardhat/'
classes = ('hardhat', 'no-hardhat', 'vest', 'machinery', 'person')
train_pipeline = [
dict(type='LoadImageFromFile'),
dict(type='LoadAnnotations', with_bbox=True),
dict(type='Resize', img_scale=(640, 640), keep_ratio=True),
dict(type='RandomFlip', flip_ratio=0.5),
dict(type='Normalize', mean=[123.675, 116.28, 103.53],
std=[58.395, 57.12, 57.375]),
dict(type='Pad', size_divisor=32),
dict(type='DefaultFormatBundle'),
dict(type='Collect', keys=['img', 'gt_bboxes', 'gt_labels']),
]
data = dict(
train=dict(
type=dataset_type,
classes=classes,
ann_file=data_root + 'annotations/train.json',
img_prefix=data_root + 'images/train/',
pipeline=train_pipeline
)
)
This is more boilerplate than Detectron2. But the payoff: I can swap in Swin-L backbone, add RepPoints head, or switch to an anchor-free paradigm by changing 3-5 lines in the config. MMDetection treats the detector as a composition of swappable modules (backbone, neck, head, loss).
Training command:
python tools/train.py configs/custom/faster_rcnn_r50_hardhat.py \
--work-dir ./work_dirs/hardhat_exp1 \
--gpu-ids 0
Faster R-CNN training took 4.9 hours for the same dataset—15% faster than Detectron2. I’m not entirely sure why, but my best guess is MMDetection’s data pipeline is more optimized (uses mmcv for image ops, which are C++-accelerated). VRAM usage was also lower: 8.2 GB vs Detectron2’s 10.1 GB at batch size 4.
The documentation is… mixed. Some pages are in English, some in Chinese, some machine-translated. I spent 40 minutes figuring out that bbox_head.loss_cls defaults to CrossEntropyLoss but can be swapped to FocalLoss via:
model = dict(
roi_head=dict(
bbox_head=dict(
loss_cls=dict(
type='FocalLoss',
use_sigmoid=True,
gamma=2.0,
alpha=0.25,
loss_weight=1.0
)
)
)
)
This level of control doesn’t exist in YOLO. In Detectron2, you’d need to subclass the ROI head.
I also tested ATSS (Adaptive Training Sample Selection), an anchor-free single-stage detector that adaptively selects positive samples based on IoU statistics during training. The key difference from YOLO is how positives are assigned: where and are the mean and standard deviation of IoU between a ground truth box and its top- candidate anchors across all pyramid levels. Samples with IoU are labeled positive. This adapts to object scale and shape better than fixed IoU thresholds.
ATSS with ResNet-50 trained in 3.8 hours (faster than Faster R-CNN since it’s single-stage) and achieved [email protected]:0.95 of 0.49—between YOLO and Faster R-CNN. Inference was 68 FPS, slower than YOLO but 3.2x faster than Faster R-CNN. A solid middle ground.

When Training Time Actually Matters (and when it doesn’t)
If you’re iterating on a proof-of-concept or your dataset is under 5,000 images, a few extra hours of training won’t kill you. I’d pick architecture flexibility over raw speed.
But if you’re training 20+ experiments with different augmentations, hyperparameters, or class balancing schemes, those hours compound. YOLO’s 2.1-hour cycle vs MMDetection Faster R-CNN’s 4.9 hours means you can run 5 experiments in the time it takes to run 2.
For production, training time is irrelevant. Inference speed, model size, and mAP are what matter. YOLOv8n at 6.3 MB and 120 FPS beats Faster R-CNN’s 170 MB and 37 FPS for most edge deployment scenarios. But if you’re running on a server with a GPU and precision matters (medical imaging, autonomous vehicles, industrial QA), the 17% mAP improvement from Cascade R-CNN justifies the larger checkpoint.
Debugging: Where YOLO Hides Information
YOLO’s training logs are clean but sparse:
Epoch GPU_mem box_loss cls_loss dfl_loss Instances Size
1/100 7.21G 1.203 2.456 1.104 142 640
You get loss values, but not per-class metrics during training. If one class is failing (my “machinery” class had 23% recall vs 81% for “hardhat”), you won’t see it until validation finishes.
Detectron2 logs are verbose:
{"loss_cls": 0.421, "loss_box_reg": 0.182, "loss_rpn_cls": 0.054,
"loss_rpn_loc": 0.039, "AP": 0.52, "AP50": 0.71, "AP75": 0.58}
MMDetection logs per-class AP during validation:
+----------+-------+-------+-------+
| class | AP | AP50 | AP75 |
+----------+-------+-------+-------+
| hardhat | 0.683 | 0.891 | 0.742 |
| no-hardh | 0.541 | 0.764 | 0.589 |
| vest | 0.598 | 0.812 | 0.651 |
| machiner | 0.387 | 0.602 | 0.412 |
| person | 0.712 | 0.902 | 0.789 |
+----------+-------+-------+-------+
This level of insight caught my “machinery” problem immediately. Turns out most machinery images were at 1920×1080 resolution, resized to 640×640, shrinking the objects below 32×32 pixels—too small for the FPN’s level (stride 8, minimum ~16px). I added Resize with img_scale=[(640, 640), (800, 800), (1024, 1024)] for multi-scale training, and machinery AP jumped to 0.51.
Preprocessing Pitfalls I Hit (so you don’t have to)
-
Color space mismatch: YOLO expects RGB. OpenCV loads BGR. If you’re using
cv2.imread()to preprocess and don’t convert, your model will train but perform 10-15% worse on real data. I lost a day to this. -
Normalization hell: YOLO auto-normalizes to [0, 1]. Detectron2 uses ImageNet mean/std by default (
mean=[103.53, 116.28, 123.675], std=[57.375, 57.12, 58.395]—these are BGR order because Detectron2 uses Caffe-style). MMDetection also uses ImageNet stats but in RGB order. If you pass pre-normalized images, you’ll double-normalize and your model will converge to garbage. -
Anchor scale mismatch: YOLO auto-calculates anchors from your dataset via k-means. Detectron2 and MMDetection use COCO-tuned anchors. If your objects are tiny (my hardhats were 40-80px), default anchors (set for 224-512px objects) won’t align well. In MMDetection, I changed
anchor_generatortoanchor_scales=[2, 4, 8](down from [8, 16, 32]) and small object AP increased 9%.
Model Checkpoints and VRAM Reality Check
| Framework | Model | Params | Checkpoint | Train VRAM (bs=4) | Inference VRAM |
|---|---|---|---|---|---|
| YOLO | YOLOv8n | 3.2M | 6.3 MB | 4.1 GB | 1.8 GB |
| YOLO | YOLOv8x | 68M | 136 MB | 16.2 GB | 5.3 GB |
| Detectron2 | Faster R-CNN R-50 | 41M | 170 MB | 10.1 GB | 3.7 GB |
| Detectron2 | Cascade R-CNN X-152 | 264M | 527 MB | 22.4 GB | 8.9 GB |
| MMDet | Faster R-CNN R-50 | 41M | 167 MB | 8.2 GB | 3.5 GB |
| MMDet | ATSS R-50 | 32M | 128 MB | 7.1 GB | 2.9 GB |
If you’re on a 16GB GPU (3080, 4070 Ti, etc.), Cascade R-CNN X-152 is out. YOLOv8x fits but leaves no headroom for larger batches. MMDetection’s ATSS is the best balance of performance and VRAM efficiency for single-GPU setups.
What I’d Pick for Different Scenarios
For a client demo in 2 days: YOLO. Ultralytics API is unbeatable for speed. Accept the mAP tradeoff.
For a research project where I’ll try 10 architectures: MMDetection. The config system makes experimentation fast once you’ve climbed the learning curve.
For production deployment on edge devices: YOLO, but export to ONNX and run via ONNX Runtime or TensorRT. I got YOLOv8n down to 4.1 MB and 210 FPS on a Jetson Orin after INT8 quantization.
For a high-stakes production system (medical, AV, industrial QA): MMDetection with Cascade R-CNN or DINO (if you have the compute). The per-class logging and modular architecture make it easier to audit and iterate.
If I need to onboard junior teammates fast: Detectron2. The API is more explicit than MMDetection’s config files, so it’s easier to read someone else’s training script and understand what’s happening.
The Part Nobody Talks About: When Your First Choice Fails
I started with YOLO because the project timeline was tight. After two weeks of hyperparameter tuning, I couldn’t get [email protected] above 0.42. My options:
- Stick with YOLO, accept lower localization precision, maybe add a post-processing NMS refinement step
- Switch frameworks, lose two weeks of work, but gain 40% better box precision
I switched to MMDetection and retrained with Cascade R-CNN. Took 3 days to port the dataset and re-run experiments. Final [email protected]: 0.63. The client was happy. I was tired.
The lesson: if you think you might need architecture flexibility later, start with MMDetection or Detectron2. Porting from YOLO is painful because the training loop, augmentation pipeline, and loss functions are all tightly coupled to Ultralytics’ design.
Curiosity I Haven’t Resolved Yet
Why is MMDetection’s training 15-20% faster than Detectron2 for the same model architecture? My best guess is mmcv‘s data pipeline, but I haven’t profiled it rigorously. If anyone has dug into this, I’m curious. Also: I haven’t tested MMDetection 3.x (released late 2023), which claims 30% faster training via a redesigned data loading system. If that holds, it’s a game-changer for large-scale experiments. Maybe I’ll revisit this when I’m not debugging YOLO memory leaks—which, by the way, are still a thing in v8.0.200.
Debugging these frameworks at 2am with a looming deadline? Keep Dark Chocolate Espresso Beans on your desk. They’ve saved more projects than I’d like to admit.
FAQ
Q: Can I train YOLO and export to ONNX for faster inference?
Yes, Ultralytics has built-in export: model.export(format='onnx'). I got 1.7x speedup on CPU inference (OpenCV DNN backend) and 1.3x on GPU (ONNX Runtime). The catch: some post-processing (NMS) stays in Python unless you use TensorRT, which fuses it into the graph.
Q: Is Detectron2 dead since development slowed?
Not dead—mature. It’s stable, well-documented, and still used in production at Meta. But if you want cutting-edge models (DINO, Grounding DINO, SAM-based detectors), you’ll need to integrate them manually or use MMDetection, which adds new models faster.
Q: Which framework has the best multi-GPU support?
MMDetection and Detectron2 both use PyTorch DDP (DistributedDataParallel) natively. YOLO added DDP in v8 but the setup is less flexible—you can’t easily customize the training loop without forking. For 4+ GPU training, I’d pick MMDetection because the config-based system makes distributed settings explicit and reproducible.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,861 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (962 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (818 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (806 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (602 views)