TLDRocket
Sign in

“Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work

The New Stack Adrian Bridgwater

Sai hit 73% on OSWorld 2.0 for Simular. The twist: it beat bigger rivals while Simular says it did it at about 2/3 the cost.

Based on reporting by The New Stack, Adrian Bridgwater — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Simular says its computer agent Sai reached a 73% success rate on OSWorld 2.0, a benchmark update released Thursday. The test covers 108 tasks that mimic long, ordinary professional work — the kind that can take skilled humans more than an hour to finish.

The company is pitching that score as more than a leaderboard flex. It says Sai landed ahead of GPT-5.6 Sol at 62.57% and Opus 5 at 70.57%, while costing about two-thirds as much as either. The point, according to co-founder and CTO Jiachen Yang, is that routine work like invoice checks, recruitment outreach, and news research should not require an expensive model built to chase headline feats.

Sai is designed for that sort of job. It runs on full desktop apps and webpages, can call APIs, and writes code. Simular says it combines frontier and specialist models, uses dedicated interfaces to see and act on a user’s computer, and is built to deliver that work at an accessible price for people and businesses. The company also leans on a neurosymbolic approach, where repeated tasks can be turned into reusable code instead of being relearned every time.

OSWorld 2.0 itself is meant to be a harder test than the earlier version. Launched on June 26 and updated on August 8 by the Executable Language Grounding Lab at the University of Hong Kong, it measures tasks in hours rather than minutes and includes messy real-world problems like hunting through email and expense reports, handling interruptions, and dealing with conflicting information. Simular says its own open-source Agent S was the first to beat the human baseline on OSWorld 1.0 last December.

There is still one awkward detail: Sai’s result is not yet on the official OSWorld 2.0 site. Simular says it is submitting the result and uploading trajectories to Hugging Face, and notes that companies often post benchmark numbers on their own sites first. That may be normal. It is also very much the modern AI version of “trust us, the spreadsheet is fine.”

My take — AI-written commentary, not fact-checked reporting

This is the right fight. Agents for invoices, bookings, and inbox sludge should be judged on cost and reliability, not on how many bragging rights they can rack up in a lab. The industry has spent long enough worshipping raw model muscle; boring work is where the bill finally comes due.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.