TLDRocket
Sign in

Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study

Together AI

Researchers introduced ReasonIF, a benchmark dataset of 300 math and science problems, to test whether large reasoning models follow user instructions during their step-by-step reasoning processes rather than just in final responses. Frontier models including GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 failed to follow reasoning instructions more than 75% of the time, with instruction-following scores dropping over 50 percentage points compared to their performance on main responses. The models showed particularly severe failures on formatting constraints like JSON formatting and uppercase-only requirements, and instruction adherence degraded further as task difficulty increased.

Why it matters

ReasonIF finds frontier LRMs fail to follow reasoning instructions >75% of the time; introduces a benchmark across languages, formatting, and length.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.