TLDRocket
Sign in

The Frontier AEO Tracker: What Astra Chooses (and every other frontier model, and what you can do about it)

Latent Space

Latent Space built a tracker of what frontier models recommend across 161 categories. It shows which products keep winning, and where model bias or search habits change answers.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Latent Space took a pile of autoresearch experiments and turned it into a Frontier AEO tracker: 6 prompt variations, 7 models with search on, and 161 categories ranging from coding agents and AI podcasts to managed databases, ASR models, Angel investors, corporate spend, and payroll software.

The setup is built around Astra, which extracts the answers and scores them with a proprietary AEO system. That score gives weight to first choices, alternatives, mentions, and even anti-recommendations, though those are described as rare. The team also pulled out the most-cited sources behind agent recommendations and a separate look at the biggest failures.

They say they checked for contamination, and that every prompt-and-answer pair is inspectable. Even so, the results show obvious biases. When models are asked for coding agent recommendations, some pairings keep recurring: Fable and Opus like Claude Code, Sol and Astra lean toward Codex and Grok, Grok likes Cursor, Muse likes Muse Code and SWE-1.7 likes Devin. The writeup also points to a few “soft biases” beyond that. GPT models recommending Claude are held up as a nice example of nonbias.

Across the full set, 28 categories have a universally dominant primary choice among all the frontier models they surveyed. That leaves a lot of room for closer contests, which is really the point of the tracker. If the answer is already stable, the market gets dull fast. If it isn’t, then the fight shifts to the categories where the models still wobble.

The more interesting comparison may be between model generations. Latent Space built separate Sol-to-Astra and Opus-to-Fable summaries to see how labs’ priorities change. One standout is source use: Sol has a median of 9 sources, Astra 5, Opus 11, and Fable 15. Astra also seems harder to shake with a light paraphrase, which the post frames as either confidence or efficiency. Either way, less randomness means AEO matters more, not less.

My take — AI-written commentary, not fact-checked reporting

This is the annoying truth about AI discovery: cleaner answers make influence games more valuable, not less. Also, the models clearly have tastes, which is a very polite way of saying they’re already doing product ranking with extra steps.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.