TLDRocket
Sign in

The State Of LLMs 2025: Progress, Problems, and Predictions

Ahead of AI Sebastian Raschka, PhD Covered by 2 sources

Large language models in 2025 focused heavily on reasoning-based post-training using reinforcement learning with verifiable rewards (RLVR) and the GRPO algorithm, pioneered by DeepSeek's R1 model released in January. DeepSeek R1 demonstrated that training state-of-the-art reasoning models costs approximately $5 million for full model development and $294,000 for post-training, significantly lower than previous estimates of $50-500 million. The industry shift toward RLVR-based reasoning models, inference-time scaling improvements, and future work on continual learning represents a fundamental change in how LLM capabilities are developed beyond just increasing model scale.

Why it matters

A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.