CLIP
CLIP is a vision-language model developed by OpenAI that aligns text and image representations in a shared embedding space, enabling zero-shot image classification without task-specific training. The model has become foundational to multimodal systems, serving as the image encoder in models like Flamingo and LLaVA, and its embeddings are used in text-to-image diffusion models such as DALL·E 2 and Stable Diffusion. Research has adapted CLIP for other languages like Chinese, integrated it into image search systems, and identified multimodal neurons within the model that activate consistently across different modalities.
Updated 3 August 2026
Specifications
No specifications recorded yet.
Latest developments
Multimodality and Large Multimodal Models (LMMs)
Chip Huyen · 2 years ago ·
47
Text-to-Image: Diffusion, Text Conditioning, Guidance, Latent Space
Eugene Yan · 3 years ago ·
14
Image search with 🤗 datasets
Hugging Face Blog · 4 years ago ·
29
Multimodal neurons in artificial neural networks
OpenAI Blog · 5 years ago ·
10
Scaling Kubernetes to 7,500 nodes
OpenAI Blog · 5 years ago ·
37
CLIP: Connecting text and images
OpenAI Blog · 5 years ago ·
21
October 2023
December 2022
November 2022
March 2022
March 2021
January 2021
Relationships
Products & technology
- OpenAI develops this model · 3 sources
- Integrated with Flamingo · 1 source
- Integrated with LLaVA · 1 source
- Chinese CLIP derived from this model · 1 source