TLDRocket
Sign in

DALL·E: Creating images from text

OpenAI

OpenAI built DALL·E, a neural net that turns written captions into images. It can dream up things that never existed—like an armchair shaped like an avocado.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has unveiled DALL·E, a neural network that takes a text caption and produces a matching image, often one that never existed until the model imagined it. The name is a mashup of Pixar's WALL-E and the surrealist painter Salvador Dalí, and that pairing tells you exactly what the researchers were going for: something between an engineer and an artist.

What makes DALL·E notable isn't just that it draws pictures from words. Plenty of systems before it could generate images conditioned on text. The trick here is generality. Feed it a weird, specific prompt — an armchair in the shape of an avocado, say — and it doesn't just retrieve something close from a training set. It composes a plausible object from scratch, blending two unrelated concepts into a coherent, often oddly convincing result.

Under the hood, DALL·E is a 12-billion-parameter version of GPT-3, trained to predict image tokens the same way GPT-3 predicts text tokens. OpenAI treats text and image data as one long stream, which lets the same transformer architecture that powers language models handle pixels too. That's the quiet architectural bet: instead of building a bespoke image-generation pipeline, extend the language-modeling trick that already worked and let it swallow a different kind of data.

The practical implications go beyond party tricks with avocado furniture. A model that can render arbitrary combinations of objects, styles, and attributes from plain English hints at tools for concept art, product mockups, and rapid prototyping — anywhere someone currently sketches an idea by hand before committing resources to it. OpenAI is positioning this as an early step, not a finished product, but the direction is clear: language models are becoming general-purpose interfaces for generating things, not just describing them.

My take — AI-written commentary, not fact-checked reporting

This is the moment text-to-image stopped being a toy and started looking like infrastructure — GPT-3's trick of predicting tokens works on pixels just as well as prose, which should worry anyone betting on narrow, bespoke AI pipelines. I'd rather see this released openly so outside researchers can poke at its biases and limits, but OpenAI teasing capability without shipping weights is becoming a predictable pattern, and a slightly annoying one.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.