DeepSeek R1
DeepSeek R1 is a reasoning-focused large language model released in January 2025 that pioneered reinforcement learning with verifiable rewards (RLVR) for post-training, achieving state-of-the-art performance on benchmarks like AIME 2024 and CodeForces while requiring significantly lower training costs ($5 million for full development, $294,000 for post-training) than previous estimates. The model's success has influenced industry-wide adoption of reasoning-based post-training approaches, and DeepSeek demonstrated that step-by-step reasoning traces from R1 can be distilled into smaller models that develop emergent reasoning abilities without reinforcement learning. R1 has become a reference model for the open-source AI community, with projects like Open R1 replicating its training pipeline and generating synthetic reasoning datasets across mathematics and other domains.
Updated 3 August 2026
Specifications
No specifications recorded yet.
Latest developments
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models
TheSequence · 1 week ago ·
43
The State Of LLMs 2025: Progress, Problems, and Predictions
Ahead of AI · 7 months ago ·
9
How to evaluate and benchmark Large Language Models (LLMs)
Together AI · 8 months ago ·
41
MedGemma: Our most capable open models for health AI development
Google DeepMind · 9 months ago ·
42
The Frontier is Open
Together AI · 1 year ago ·
30
The NLP Course is becoming the LLM Course
Hugging Face Blog · 1 year ago ·
19
QwQ-32B: Embracing the Power of Reinforcement Learning
Qwen · 1 year ago ·
38
Open R1: Update #2
Hugging Face Blog · 1 year ago ·
18
July 2026
December 2025
Researcher compiles curated list of LLM research papers from July-December 2025 Research publication
November 2025
October 2025
Google DeepMind releases multiple new Gemma model variants including Gemma 3, MedGemma, and T5Gemma Model release
June 2025
April 2025
March 2025
February 2025
January 2025
Hugging Face releases tutorials and open-source projects replicating DeepSeek R1's reinforcement learning training methodology Open source release
Relationships
Products & technology
- DeepSeek develops this model · 4 sources
- Derived from DeepSeek 32B · 1 source
- Derived from DeepSeek 7B · 1 source
- Open-R1 derived from this model · 1 source
- OpenR1-Math-220k derived from this model · 1 source
- Hugging Face integrated with this model · 1 source