TLDRocket
Sign in

GPT-4

OpenAI Covered by 2 sources

OpenAI just launched GPT-4, a new AI that can read both text and images and write text back. It aces professional exams but still trips on everyday reasoning.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI dropped GPT-4 this week, calling it the next big step in its long bet on scaling up deep learning. Unlike its predecessor, this model takes both images and text as input, though it still only talks back in words. That's a meaningful shift: the company is no longer just building better text predictors, it's building systems that can look at a photo, a chart, or a screenshot and reason about what they see.

What's striking is how OpenAI frames GPT-4's abilities. The company says it performs at a human level on a range of professional and academic benchmarks, the kind of standardized tests that separate lawyers from paralegals or med students from doctors. That's a bold claim, and it's backed by numbers rather than vibes, which is more than can be said for most AI marketing copy.

But OpenAI is also upfront that the model is far from human in a lot of ordinary situations. Passing a bar exam is one thing, navigating the messy, contradictory, poorly specified problems of real life is another. The gap between benchmark performance and real-world usefulness has bitten every big model release so far, and there's no reason to think GPT-4 magically closes it.

Still, the multimodal piece matters more than the benchmark scores, longer term. Text-only models hit a ceiling in how much of the world they can actually engage with. A model that can parse a diagram, a handwritten note, or a screenshot of buggy code opens up a much wider set of tasks, from tutoring to accessibility tools to plain old customer support. GPT-4 won't be the last model to add vision, but it's the one that made OpenAI's rivals scramble to catch up.

My take — AI-written commentary, not fact-checked reporting

I'll believe the human-level claims when independent researchers, not OpenAI's own paper, run the tests on problems that weren't lurking in the training data. Benchmark scores are cheap theater compared to what actually ships in products, and OpenAI still won't tell us what's in the training set or how big the thing is. Closed weights, closed data, open marketing, that's the pattern, and it's exhausting.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.