Basecamp Bench
TLDR Dev ● Covered by 7 sources
A benchmark tested five AI models (Anthropic's Fable 5, OpenAI's GPT-5.6 Sol and GPT-5.5, SpaceX's Grok 4.5, and Google's Gemini Pro 3.1) by having them build a frontend and backend for a Basecamp project from scratch. Fable 5 scored highest on both frontend and backend tracks, while Grok 4.5 completed both builds in 37 minutes for $9.30, offering the best speed-to-cost ratio despite visible polish gaps. The results show significant performance variation across models, with frontend work revealing larger gaps than backend work, suggesting that achieving high-quality UI polish remains a differentiator between leading AI models.
Why it matters
Recent testing compared the performance of GPT-5.6 Sol, Fable 5, and Grok 4.5 in solving a complex software problem, showing Fable 5 as the top performer in both frontend and backend development.
Also covered by
- TheSequence — The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
- Platformer — OpenAI's big launch — and bigger departure
- The Neuron — GPT-5.6 Sol Live Testing
- Deep Learning Weekly — Deep Learning Weekly: Issue 463
- Zvi (Don't Worry About the Vase) — AI #176 Part 1: Doing It Live
- TLDR Dev — GLM 5.2 and the coming AI margin collapse (part 1)