Alibaba President: AI agents can talk, but can they actually do the work?
Fortune Kuo Zhang ● Covered by 2 sources
Alibaba’s president says AI agents need tests for real work, not just smart answers. That matters because in commerce, a good-sounding mistake can still wreck an order.
Based on reporting by Fortune, Kuo Zhang — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Joshua Stancle, who runs Clean Saint in Los Angeles, says AI now helps him with sourcing, marketing, web development, and customer support. His line was memorable: “In a sense, there are ten of me.” But that only works if the extra nine can actually finish the job, not just talk well about it.
That’s the point Alibaba.com is pushing. The company says the real question is not which model is smartest in the abstract, but which system can handle the ugly details of commerce: supplier calls, customs forms, shipping problems, returns, and all the small failures that happen between a request and a completed order. A model can reason. A business needs execution.
So Alibaba’s Accio team built CommerceAgentBench, an open-source benchmark on GitHub that grades outcomes instead of chatty answers. It includes 107 end-to-end tasks drawn from procurement, logistics, product listing, fulfillment, and after-sales service. Alibaba says those tasks were assembled from its own data: 10 million active small-business users, 1.6 million real conversations, and 200,000 execution traces, sorted into seven categories of commercial work.
The results were useful and a little sobering. The strongest frontier model tested completed 61.7% of the tasks. That leaves a lot of room for errors, and the failures clustered in places that look painfully familiar: payment anomalies buried in long supplier threads, landed-cost calculations with too many moving parts, after-sales disputes where documents conflict, and multi-leg shipping routes that simply break down.
And the leaderboard wasn’t settled either. No single model won across the board. One model led on request-for-quote work and market research, then fell behind on claims settlement and listing compliance, where another model came out ahead. A third was best at publishing products and handling returns. That is the real message here: model choice depends on the workflow, not on a braggy ranking chart.
Alibaba’s argument is basically for precision delegation. Hand over the routine parts where the pass rate is high. Keep a person on the weird compliance questions, the tricky negotiations, and the exceptions that make commerce messy. The company says other fields will need the same kind of outcome-based tests, and that those benchmarks should be open. Hard to argue with that, unless the plan was to keep pretending a fluent bot is the same thing as a competent operator.
My take — AI-written commentary, not fact-checked reporting
This is the right fight to pick. The industry still loves scoring models like they’re debating chess grandmasters, when businesses need someone who can keep a shipment from dying in Ningbo. Open benchmarks beat glossy demos every time, because the bill always comes due in the workflow, not the keynote.
Read more about this at: Fortune