Tag: vision-language models
-
CLIP ViT-L/14 Zero-Shot ImageNet Accuracy: 76.2% Without Fine-Tuning
CLIP ViT-L/14@336px hits 76.2% zero-shot ImageNet top-1 accuracy. We break down how contrastive pretraining works and where it fails.
CLIP ViT-L/14@336px hits 76.2% zero-shot ImageNet top-1 accuracy. We break down how contrastive pretraining works and where it fails.