TLDRocket
Sign in

The State Of LLMs 2025: Progress, Problems, and Predictions

Ahead of AI Sebastian Raschka, PhD Covered by 2 sources

Large language models in 2025 focused heavily on reasoning-based post-training using reinforcement learning with verifiable rewards (RLVR) and the GRPO algorithm, pioneered by DeepSeek's R1 model released in January. DeepSeek R1 demonstrated that training state-of-the-art reasoning models costs approximately $5 million for full model development and $294,000 for post-training, significantly lower than previous estimates of $50-500 million. The industry shift toward RLVR-based reasoning models, inference-time scaling improvements, and future work on continual learning represents a fundamental change in how LLM capabilities are developed beyond just increasing model scale.

Why it matters

A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.