TLDRocket
Sign in

Advanced Reasoning

49 summarised stories about Advanced Reasoning, each linking back to the original source. Browse all topics →

Thursday, 1 May 2025

Why We Think

Lilian Weng 1 year ago

A research post reviews how test-time compute and chain-of-thought reasoning improve language model performance by enabling models to spend more computation on problem-solving. Key developments include reinforcement learning approaches like o1 and o3 that use policy gradient algorithms, parallel sampling methods like best-of-N and beam search with process reward models, and sequential revision techniques, with recent work showing optimal ratios of sequential to parallel compute depending on problem difficulty. The effectiveness stems from treating computation as a learnable resource, drawing analogies to human System 2 thinking, and using latent variable modeling to optimize over intermediate reasoning steps.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.