CLIP: Connecting text and images
OpenAI 5 years ago 22
OpenAI introduced CLIP, a neural network that learns visual concepts by connecting text descriptions with images. CLIP performs zero-shot classification on visual benchmarks without requiring task-specific training data. This approach enables the model to recognize new image categories by simply receiving their text names, reducing the need for large labeled datasets in computer vision tasks.