TLDRocket
Sign in

Model Evaluation

56 summarised stories about Model Evaluation, each linking back to the original source. Browse all topics →

+ Follow this topic

Tuesday, 18 August 2026

When a model reads a drug's class from its name—not its knowledge

Allen Institute (AI2) 2 weeks ago 21

Researchers found that Olmo 3 often relies on drug name affixes like -pril and -olol rather than actual drug knowledge when answering health questions, with 51–59% of tested drugs showing little sign of specific knowledge and 12–18% appearing purely affix-driven. The team used diagnostic tests swapping drug name components with nonsense words and traced the behavior to training data, finding that rarer drugs in the corpus triggered more affix-based inference. This shortcuts approach to drug information matters for health advice because while affixes do encode real pharmacological classes, relying on them instead of actual drug knowledge risks poor medical guidance.

DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AI 2 weeks ago 19 2 sources

Researchers compared DeepSeek V4 Pro 0813 and GPT-5.6 Sol on DeepSWE, a software engineering benchmark with 113 tasks across multiple languages and domains. Pro costs $0.24 per rollout while Sol costs $8.37—a 35x difference—with Sol achieving 72.7% pass@1 versus Pro's 62.8%, but Pro reaching 88.5% pass@4 versus Sol's 85.8%. A cascading approach that runs Pro first and escalates to Sol only on failures solves 83.0% of tasks for $3.35 each, outperforming either model alone and beating a perfect oracle router.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.