TLDRocket
Sign in

Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study

Together AI

Researchers introduced ReasonIF, a benchmark dataset of 300 math and science problems, to test whether large reasoning models follow user instructions during their step-by-step reasoning processes rather than just in final responses. Frontier models including GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 failed to follow reasoning instructions more than 75% of the time, with instruction-following scores dropping over 50 percentage points compared to their performance on main responses. The models showed particularly severe failures on formatting constraints like JSON formatting and uppercase-only requirements, and instruction adherence degraded further as task difficulty increased.

Why it matters

ReasonIF finds frontier LRMs fail to follow reasoning instructions >75% of the time; introduces a benchmark across languages, formatting, and length.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.