TLDRocket
Sign in

Addendum to GPT-4o System Card: 4o image generation

OpenAI Covered by 2 sources

OpenAI baked image generation directly into GPT-4o instead of bolting on DALL·E 3. It's way better at photorealism and can actually edit images you feed it, not just make new ones.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has quietly rewritten how ChatGPT handles pictures. The new system, described in an addendum to the GPT-4o system card, ditches the old DALL·E 3 pipeline in favor of native image generation built right into the 4o model itself. That's a meaningful architectural shift, not just a version bump.

The practical upshot is a jump in photorealism. Earlier DALL·E-based outputs often had that telltale AI sheen — slightly off hands, waxy skin, compositions that looked generated rather than shot. OpenAI says 4o's image outputs close that gap substantially, producing results that read as genuinely photographic rather than illustrative.

The bigger functional change is input handling. Previous models mostly took a text prompt and spat out a picture. 4o can take an actual image as input and transform it, meaning users can now edit, restyle, or manipulate existing photos through the same model that handles text and reasoning, rather than routing through a bolted-on generator. That folds image editing into the same conversational loop as everything else ChatGPT does.

OpenAI frames this as enough of a capability jump to warrant its own safety addendum, on top of the existing GPT-4o system card. Anytime a company publishes a dedicated safety document for one feature, it's a decent signal they expect that feature to get used — and misused — a lot more than what came before it.

My take — AI-written commentary, not fact-checked reporting

Folding image generation directly into the base model instead of duct-taping DALL·E on the side is the correct move, and honestly overdue — separate specialist models bolted onto a chat interface always felt like a stopgap. The photorealism jump is the real story here, and it's the part that should worry people more than any chatbot text output ever has, because a convincing fake photo travels faster and gets believed harder than a paragraph of AI-generated nonsense.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.