TLDRocket
Sign in

Fable 5 Vs Opus 4.8: Outcomes-Based Assessments Are A Massive Warning For Frontier AI Labs

TLDR Dev

The author tested Fable 5 and Opus 4.8 on a real-world task of rebuilding a website to improve conversion rates and found both models scored 0 on outcomes-based metrics despite producing functional technical artifacts. Both models failed to include basic features like conversion tracking and security without explicit prompting, and a smaller open-source model (Gemma 4) performed equally well at zero token cost when given sufficient context. Companies like Eli Lilly are moving away from expensive frontier AI models toward smaller purpose-built models fine-tuned on proprietary data, signaling that enterprise customers prioritize outcomes-based value over frontier model capabilities.

Why it matters

Recent assessments show that advanced AI models, specifically Fable 5 and Opus 4.8, fail to meet real-world business outcomes, scoring zero in improving website conversion rates despite appearing capable on artifact metrics, with companies opting for smaller purpose-built models.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.