TLDRocket
Sign in

Generalized Visual Language Models

Lilian Weng

Researchers have developed multiple approaches to extend pre-trained language models to process visual information alongside text, grouping these vision-language models (VLMs) into four categories: joint training with image-text embeddings, frozen language model prefixes, cross-attention fusion mechanisms, and combined models without training. Notable models include VisualBERT trained on MS COCO with masked language modeling and sentence-image prediction objectives, SimVLM mixing 4,096 image-text pairs with 512 text-only documents per batch, and CM3 trained on close to 1 trillion tokens of web data tokenized to 256 tokens per image. These approaches enable language models to perform vision-language tasks like image captioning and visual question-answering while preserving or leveraging existing pre-trained linguistic capabilities.

Why it matters

Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection network as a vision encoder to capture visual features and then produce text via a text decoder. Given a large amount of existing literature, in this post, I would like to only focus on one approach for solving vision language tasks, which is to extend pre-trained generalized language models to be capable of consuming visual signals.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.