It cost $33 to build a virtual Union Square. Here’s what the agents got wrong.
The New Stack Amanda Caswell
PhiloLabs built a fake Union Square with AI agents for about $33. It worked in the browser, but the screenshots exposed what the code missed.
Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
PhiloLabs tried a simple but nasty question: can AI coding agents build something that looks right, not just something that runs? The answer, at least for a virtual San Francisco Union Square, was yes — and no. In about two hours, Claude Fable 5.1 agents produced a working Three.js version of the square in the browser, using real geographic data and reference images.
The scene was not toy-sized. It included 453 building footprints, 75 custom façades, 129 named storefronts, 220 pedestrians and 109 vehicles, with Powell Street cable cars moving through the space. The whole run burned through roughly 8 million tokens and came out to about $33 in API calls. That’s not nothing, but it is cheap enough to make the experiment feel less like a demo and more like a test of process.
PhiloLabs broke the work into subagents. Some handled geographic research, others geometry, textures, storefronts and the rest of the scene. Then Playwright entered the loop, marching through 34 fixed camera positions and taking screenshots that could be compared with photos of the real Union Square. That produced 147 comparison sheets, which made visual mistakes easier to spot — the kind where a building is technically in the right place but still looks wrong, or a storefront ends up on the wrong side of the street.
The company also had specialist agents review the build, with separate passes for architecture, geography, technical art and interactions. Those reviewers generated nine reports, which became the punch list for the next round. It sounds tidy, but there’s a catch: if the source data is incomplete, another agent can miss the same gap the first one did. A building can be placed from OpenStreetMap data and USGS elevation data, yet still need the model to invent a façade where the photos don’t show enough.
That’s the real lesson here. The hard part is not getting agents to spit out code; it’s getting them to make visual judgments that hold up when the inputs are patchy and the scene has to feel believable from more than one angle. Screenshots help. They just don’t magically turn approximation into truth.
My take — AI-written commentary, not fact-checked reporting
This is the sort of experiment that matters because it’s annoyingly honest. Agents are getting decent at assembling the scaffolding, but they still need adults in the room when the job turns into taste, judgment, and missing data. The hype machine loves “it runs”; users care about “it looks wrong.”
Read more about this at: The New Stack
Related stories
Opus Outshines Even Fable, Inside the Hugging Face Hack, AI Companies Spend Big for Compute
The Batch ·
41
😺Anthropic launched Fable 5.1: and now, the agents cost less
The Neuron · 1 day ago ·
5
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
MarkTechPost · 2 days ago ·
6