TLDRocket
Sign in

Multimodality and Large Multimodal Models (LMMs)

Chip Huyen

Large Multimodal Models (LMMs) combine language models with the ability to process multiple data types like text, images, and audio, with CLIP and Flamingo serving as foundational examples demonstrating how to align different modalities into shared embedding spaces. CLIP achieved competitive zero-shot performance on image classification tasks and has been adopted as the image encoder in systems like Flamingo and LLaVA. The shift toward multimodal systems enables practical applications across healthcare, robotics, e-commerce, and accessibility use cases that require processing mixed data types.

Why it matters

For a long time, each ML model operated in one data mode – text (translation, language modeling), image (object detection, image classification), or audio (speech recognition). However, natural intelligence is not limited to just a single modality. Humans can read, talk, and see. We listen to music to relax and watch out for strange noises to detect danger. Being able to work with multimodal data is essential for us or any AI to operate in the real world. OpenAI noted in their GPT-4V system card that “incorporating additional modalities (such as image inputs) into LLMs is viewed by some as a key frontier in AI research and development.” Incorporating additional modalities to LLMs (Large Language Models) creates LMMs (Large Multimodal Models). Not all multimodal systems are LMMs. For example, text-to-image models like Midjourney, Stable Diffusion, and Dall-E are multimodal but don’t have a language model component. Multimodal can mean one or more of the following: Input and output are o

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.