Computer use is far from solved
Steelman Labs 1 month ago 48
Agentic AI models show only 20.6% task completion on the OSWorld-V2 benchmark despite claims of 60-97% performance on earlier benchmarks, revealing that real-world computer use remains far from solved. The gap exists because models often bypass user interfaces entirely through API calls or scripting rather than performing basic UI actions like clicking buttons, which humans do effortlessly. Steelman Labs argues the core problem is architecture: frontier models waste computation on perception and clicking instead of reasoning, and solving this requires separating planning from execution with a dedicated motor-control system rather than scaling up larger models.