TLDRocket
Sign in

Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial

Hugging Face Covered by 2 sources

A tutorial demonstrates how to recreate DeepSeek R1's "aha moment"—where a model learns to allocate more thinking time to problems through reinforcement learning—using Group Relative Policy Optimization (GRPO) on the Countdown Game puzzle task. The training setup uses a Qwen 2.5 3B model on 4 NVIDIA H100 GPUs, with full training completing in approximately 6 hours across 450 steps. By step 450, the model achieves 50% success rate in solving countdown equations and spontaneously shifts from word-based reasoning to programmatic trial-and-error approaches without explicit instruction.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.