TLDRocket
Sign in

Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less

Together AI

Together AI pit the open Kimi K2.7 Code model against Claude Fable 5 on building 12 landing pages. Kimi cost 94% less and scored nearly as well — open-source is closing the gap fast.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Together AI just published a fairly blunt experiment: build the same 12 landing pages with Kimi K2.7 Code, an open-source model, and with Claude Fable 5, and see what breaks down. The headline number is stark. Kimi came in roughly 94% cheaper on average, about 16 times less than Fable and 8 times less than Claude Opus 4.8, while landing within a few points of Fable's quality score on nearly every page.

The test wasn't just "type a prompt, get a page." Together first tried that baseline approach across categories like a B2B SaaS tool, a rooftop speakeasy bar, and a SQL-to-charts dev tool. Both models produced pages that looked, in the team's words, recognizably AI-generated — competent but generic. The real shift came when they hooked Kimi up to a custom MCP server stocked with screenshots of well-designed landing pages and UI components. Because Kimi K2.7 Code is multimodal, it could actually look at those reference images alongside the text prompt, not just read a description of good design. The before-and-after on the speakeasy page tells the story: sharper hierarchy, better typography, fewer broken-image placeholders, faster load times.

Cost is where this gets concrete rather than theoretical. One B2B SaaS landing page cost 4 cents to generate with Kimi. The identical prompt cost $1.09 with Fable — nearly 27 times more. Scale that to a realistic workflow, where you're generating dozens of variants to find a direction worth refining, and the arithmetic adds up fast: 100 pages with Kimi versus Fable saves around $94.

To judge quality without just eyeballing screenshots, Together fed GPT-5.5 a rubric covering positioning, visual direction, content structure, craft, responsiveness, and technical execution, then had it score each page from 0 to 100. Fable won on both sample pages shown, but the margin was narrow enough that Together called the tradeoff reasonable given the price gap.

The bigger takeaway buried in the writeup: prompts alone don't get you far with either model. What actually moved quality was giving Kimi visual context to work from, not swapping in a pricier model. That's a useful reminder that a lot of the perceived gap between open and closed models is really a gap in how well people are feeding them information.

My take — AI-written commentary, not fact-checked reporting

This is exactly the kind of result that should worry anyone charging premium prices for closed models on commodity tasks like landing pages. The lesson isn't "open models caught up" — it's that most of the quality gap people attribute to model capability is actually a context and tooling problem, and once you fix that with something as simple as an MCP server full of screenshots, a 94% cost cut for a few percentage points of score is not a close call. I'd bet this pattern shows up in dozens of other coding-agent workflows once people bother to test it properly instead of assuming premium equals better.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.