TLDRocket
Sign in

GPT 5.6 Sol is the best "vision" model OpenAI ever released

Roboflow Blog

OpenAI's new GPT-5.6 lineup (Sol, Terra, Luna) makes a huge leap in vision tasks like object detection and counting. Sol still trails Gemini 3.5 Flash on price and top scores, but the gap OpenAI needed to close just got a lot smaller.

Based on reporting by Roboflow Blog — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI dropped GPT-5.6 last week with three flavors — Sol, Terra, and Luna — and leaned hard on computer-use demos during the release stream, showing off agents clicking around desktop apps and parsing 3D scenes. Both of those tricks live or die on how well a model actually sees. So Roboflow ran the new lineup through its upcoming VLM benchmark, which tests detection, counting, OCR, and data extraction, and the results tell a clear story: Sol is the best vision model OpenAI has shipped.

The biggest jump is in object detection. GPT-5.5 scored a rough 13.8 mAP@50 in Roboflow's tests, essentially unusable for real detection work. Sol jumped to 46.2, with Terra and Luna close behind at 44.7 and 43.3. That takes detection from a glaring weakness to something people can actually build on. Sol also handled document layout well, picking out titles, tables, and signatures, and it managed dense, cluttered scenes like piles of pills or eggs where VLMs typically choke because every object has to be described as text, coordinates and all.

Counting improved across the board too. Sol hit 73.0%, up from GPT-5.5's 64.9%, while even Luna, the cheapest of the three, beat the old baseline at 66.2%. Sol handled trickier counting logic as well, like tallying overlapping metal brackets or counting bullet holes only within specific target zones. It still stumbled on things like blister packs with hard-to-distinguish filled and empty slots, and one candy example where it's unclear if it miscounted or misread the category entirely.

OCR didn't move much — Sol scored 90.7% versus GPT-5.5's 91.2%, and actually fell behind on targeted text extraction, 82.5% versus 87.6%. It still nailed some hard cases, like reading a tire size printed on a worn, curved tire surface, or pulling a live score off a hockey broadcast in the right format. But it flubbed something as simple as a small, low-contrast expiration date on a blister pack.

None of this comes free. Sol runs close to 10 seconds per image and costs about 2.5 cents each, making it the second-priciest model behind Claude Fable 5. Gemini 3.5 Flash still beats Sol on both detection and counting while costing less, around 0.8 cents per image, which keeps it the more practical pick for anyone processing images at scale. OpenAI also confirmed to Roboflow that Sol gets shaky on very large images unless you crank up reasoning effort, which then costs more time and money — so resizing images before sending them is the real-world fix for now.

Still, the direction is obvious. OpenAI has clearly decided vision matters, and Sol proves it, even if Gemini keeps the crown on raw cost-performance for detection and counting.

My take — AI-written commentary, not fact-checked reporting

Sol closing the gap with Gemini 3.5 Flash matters more than the headline benchmark wins, because OpenAI spent years treating vision as an afterthought bolted onto a language model. The instability on large images and the higher cost per image show this is still a work in progress, not a finished product, and anyone building detection-heavy pipelines should keep comparing options rather than assuming the biggest name wins by default.

Read more about this at: Roboflow Blog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.