TLDRocket
Sign in

Models & Research

1194 summarised stories in Models & Research, each linking back to the original source. Browse all topics →

Tuesday, 1 September 2026

Claude Fable 5.1 made me a really nice animated pelican

Simon Willison’s Weblog 2 days ago 5 5 sources

Anthropic released Claude Fable 5.1 and the article tests it by generating an animated SVG of a pelican riding a bicycle. Terminal-Bench-Science 0.1 is reported at a 52.6% score for Fable 5.1. Higher reasoning-effort settings produce much longer outputs and a better-quality pelican SVG, enabling a separate run that animates the benchmark “pelican” without paying the full cost again.

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face 2 days ago 35 2 sources

BenchMIRT was introduced as a method to audit LLM benchmarks at the level of individual prompts, estimating which underlying capabilities each question measures. BenchMIRT was trained on benchmarking results from 100 LLMs across more than 34K questions. It recovered two stable capability dimensions—safety and general reasoning—showing that averaging benchmark items can hide mixed signals and that using only 10% to 50% of questions often preserves the same capability picture while predicting held-out answers 79% of the time.

Introducing agentic video understanding with Gemini

Google 3 days ago 12 6 sources

Google DeepMind launched agentic video understanding for video analysis in Gemini models, using goal-directed video tools instead of fixed-rate “static” processing. It cuts token consumption by up to 88% and analysis costs by up to 66%, while improving accuracy by up to 7%. The feature is now available via the Gemini API (set processing to agentic) and will roll out in the Gemini app and power YouTube’s Ask YouTube.

Meta just beat OpenAI and Google at real-time transcription

The New Stack 3 days ago 23 8 sources

Meta’s Superintelligence Labs launched Muse Voice Transcribe, a real-time speech recognition model that performs ahead of comparable streaming competitors on reported benchmarks. Its AA-WER Streaming word error rate is 3.1%, versus Cartesia Ink-2 at 3.4%, ElevenLabs’ Scribe v2 Realtime at 3.6%, GPT Live Transcribe at 3.9%, and Gemini 3.5 Transcribe Live at 4.0%. The model becomes available via the Meta Model API, Meta AI for Mac, and Muse Code, while Meta keeps it closed-weight rather than releasing open weights.

Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance

MarkTechPost 3 days ago 17

Researchers from Princeton University, Ant Group, and Stanford University introduced AQuA, a two-part framework where language-model agents iteratively improve quantitative research while the evaluation setup stays fixed to prevent evidence leakage. On a crypto five-minute universe across 20 research epochs, AQuA’s combined validation Spearman IC rose to about 0.190 (vs 0.171 for AlphaMemo and lower for several baselines). The approach changes research by making leakage-inducing actions unavailable and by having agents emit only constrained factor expressions or configuration diffs, yielding a positive equity long/short book with about +2.50 Sharpe after a volatility-targeting overlay.

After raising $10M in funding, Visko debuts Orbis, its first live model for generating long-form videos

SiliconANGLE 3 days ago 42

Visko Platform Inc. raised $10 million in pre-seed funding and debuted Orbis, its first live foundation model for generating long-form, continuously running interactive video worlds. Orbis streams at 4K resolution and 24 frames per second, and Visko claims it can generate worlds for hours without observable degradation. The model shifts AI video generation from stopping-and-restarting short offline clips to continuous generation with real-time user interventions that update without breaking coherence.

The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

TheSequence 3 days ago 14

PrismML released Bonsai 27B in July 2026, a multimodal distilled model with long context designed to run on-device. The binary version is about 3.9 GB, intended to fit in the memory budget of a high-end phone. The release shifts how distillation is framed toward end-to-end low-bit training and quantization, treating model “lineage” and packaged descendants as the meaningful unit rather than the individual checkpoint.

BenchMIRT: What are LLM benchmarks actually measuring?

Allen Institute (AI2) 3 days ago 30 2 sources

BenchMIRT audits large language model benchmark scores at the level of individual prompts by using multidimensional item-response theory to estimate the capabilities each task tests. It was trained on results from 100 LLMs covering 16 benchmarks and more than 34,000 questions. Benchmark scoring becomes more interpretable by separating signals like safety and general reasoning, revealing cases where “safety” benchmarks are driven largely by reasoning or other mixed effects, and enabling smaller question sets while largely preserving the underlying capability picture.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.