Anthropic released Claude Fable 5.1 and the article tests it by generating an animated SVG of a pelican riding a bicycle. Terminal-Bench-Science 0.1 is reported at a 52.6% score for Fable 5.1. Higher reasoning-effort settings produce much longer outputs and a better-quality pelican SVG, enabling a separate run that animates the benchmark “pelican” without paying the full cost again.
BenchMIRT was introduced as a method to audit LLM benchmarks at the level of individual prompts, estimating which underlying capabilities each question measures. BenchMIRT was trained on benchmarking results from 100 LLMs across more than 34K questions. It recovered two stable capability dimensions—safety and general reasoning—showing that averaging benchmark items can hide mixed signals and that using only 10% to 50% of questions often preserves the same capability picture while predicting held-out answers 79% of the time.
Google DeepMind launched agentic video understanding for video analysis in Gemini models, using goal-directed video tools instead of fixed-rate “static” processing. It cuts token consumption by up to 88% and analysis costs by up to 66%, while improving accuracy by up to 7%. The feature is now available via the Gemini API (set processing to agentic) and will roll out in the Gemini app and power YouTube’s Ask YouTube.
Meta’s Superintelligence Labs launched Muse Voice Transcribe, a real-time speech recognition model that performs ahead of comparable streaming competitors on reported benchmarks. Its AA-WER Streaming word error rate is 3.1%, versus Cartesia Ink-2 at 3.4%, ElevenLabs’ Scribe v2 Realtime at 3.6%, GPT Live Transcribe at 3.9%, and Gemini 3.5 Transcribe Live at 4.0%. The model becomes available via the Meta Model API, Meta AI for Mac, and Muse Code, while Meta keeps it closed-weight rather than releasing open weights.
Researchers from Princeton University, Ant Group, and Stanford University introduced AQuA, a two-part framework where language-model agents iteratively improve quantitative research while the evaluation setup stays fixed to prevent evidence leakage. On a crypto five-minute universe across 20 research epochs, AQuA’s combined validation Spearman IC rose to about 0.190 (vs 0.171 for AlphaMemo and lower for several baselines). The approach changes research by making leakage-inducing actions unavailable and by having agents emit only constrained factor expressions or configuration diffs, yielding a positive equity long/short book with about +2.50 Sharpe after a volatility-targeting overlay.
Visko Platform Inc. raised $10 million in pre-seed funding and debuted Orbis, its first live foundation model for generating long-form, continuously running interactive video worlds. Orbis streams at 4K resolution and 24 frames per second, and Visko claims it can generate worlds for hours without observable degradation. The model shifts AI video generation from stopping-and-restarting short offline clips to continuous generation with real-time user interventions that update without breaking coherence.
PrismML released Bonsai 27B in July 2026, a multimodal distilled model with long context designed to run on-device. The binary version is about 3.9 GB, intended to fit in the memory budget of a high-end phone. The release shifts how distillation is framed toward end-to-end low-bit training and quantization, treating model “lineage” and packaged descendants as the meaningful unit rather than the individual checkpoint.
BenchMIRT audits large language model benchmark scores at the level of individual prompts by using multidimensional item-response theory to estimate the capabilities each task tests. It was trained on results from 100 LLMs covering 16 benchmarks and more than 34,000 questions. Benchmark scoring becomes more interpretable by separating signals like safety and general reasoning, revealing cases where “safety” benchmarks are driven largely by reasoning or other mixed effects, and enabling smaller question sets while largely preserving the underlying capability picture.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.