TLDRocket
Sign in

Basecamp Bench

Stephen M. Walker II Covered by 7 sources

Someone built a real benchmark to test GPT-5.6 Sol, Claude Fable 5, and Grok 4.5 by making them clone Basecamp from scratch. Fable 5 crushed the competition on quality, but Grok 4.5 did the job in 37 minutes for under ten bucks.

Based on reporting by Stephen M. Walker II — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Every AI lab loves to claim their newest model is a frontend wizard, so one developer decided to actually check by making three flagship models build the same app: a Basecamp clone, judged on 20 different dimensions instead of a single vibes-based score. The results, published as Basecamp Bench, complicate the marketing copy quite a bit.

Fable 5, Anthropic's priciest model, won both the frontend and backend tracks outright, and by a wide margin. It matched the real Basecamp interface within a few percentage points, needing only minor polish a human designer could knock out in an afternoon. GPT-5.6 Sol, despite OpenAI's claims of improved frontend chops, paired a genuinely disciplined backend with a frontend that looked fine at a glance but felt shallow underneath. Grok 4.5 was the value play: it finished both builds in 37 minutes for $9.30, nearly matching Sol on backend depth, though its interface had real layout problems.

The more interesting finding is where these models actually diverge. Backend scores clustered tightly, since most models can find the right routes; what separated them was whether they enforced invariants and failed honestly instead of faking success. Frontend scores spread out much more, because getting basic functionality on screen is easy but nailing spacing, icons, and transitions that make something feel finished is brutally hard. The author leaned on an old programming maxim here: the last 10% of a project eats as much effort as the first 90%, and climbing from an 8 to a 9 on this benchmark takes far more work than climbing from a 5 to a 6.

Running Sonnet 5 and GPT-5.6 Sol five extra times each revealed something benchmark screenshots usually hide: run-to-run variance. Sol's best attempt beat Sonnet's worst, and vice versa, meaning two people forming opinions about either model could easily be looking at completely different tiers of output. Not every model made it into the results at all — GLM 5.2 couldn't complete the single-file constraint no matter how the harness was adjusted, Gemini Flash 3.5 kept losing task state and crashing tool calls, and Gemini Pro 3.1 finished but scored barely above 3 out of 10, apparently unable to sustain focus across the benchmark's long time horizon.

Building just the GPT-5.6 Sol submission burned through 682 million tokens and $449 in model costs, which says something about what real agentic software work actually costs versus a quick chat completion. The full runner, prompts, rubrics, and reference material are open on GitHub, and the author plans to slot in Gemini 3.5 Pro and, if the rumors hold, GPT-6 later this summer.

My take — AI-written commentary, not fact-checked reporting

This is the kind of benchmark I actually trust: one messy, realistic task dissected twenty ways instead of a leaderboard number designed to flatter whoever paid for the eval. The fact that Fable 5's premium price bought genuinely superior craftsmanship, while Grok 4.5 proved cheap models can get you 90% there for pocket change, is exactly the tradeoff open competition should keep sharpening — and it's a reminder that a lot of frontier marketing about 'frontend mastery' is still mostly vibes until someone runs the numbers five times and shows you the variance.

Read more about this at: Stephen M. Walker II

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.