D
Demerzel
Video
Contrastive Language-Image Pretraining (CLIP)
The shared text-image embedding space behind modern multimodal systems: contrastive InfoNCE training on image-caption pairs and zero-shot classification.