TLDRocket
Sign in

Generalized Visual Language Models

Lil'Log

Researchers have developed multiple approaches to extend pre-trained language models to process visual information alongside text, grouping these vision-language models (VLMs) into four categories: joint training with image-text embeddings, frozen language model prefixes, cross-attention fusion mechanisms, and combined models without training. Notable models include VisualBERT trained on MS COCO with masked language modeling and sentence-image prediction objectives, SimVLM mixing 4,096 image-text pairs with 512 text-only documents per batch, and CM3 trained on close to 1 trillion tokens of web data tokenized to 256 tokens per image. These approaches enable language models to perform vision-language tasks like image captioning and visual question-answering while preserving or leveraging existing pre-trained linguistic capabilities.

Why it matters

Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection network as a vision encoder to capture visual features and then produce text via a text decoder. Given a large amount of existing literature, in this post, I would like to only focus on one approach for solving vision language tasks, which is to extend pre-trained generalized language models to be capable of consuming visual signals.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.