Reasoning models struggle to control their chains of thought, and that’s good
OpenAI Blog
OpenAI introduced CoT-Control, a method to test whether reasoning models can direct their internal thought processes when prompted. The researchers found that current reasoning models fail to reliably follow instructions about how to structure their reasoning, even when explicitly asked to do so. This limitation suggests that monitoring a model's reasoning chains could serve as a safety mechanism, since the model cannot easily manipulate its own thinking process on command.
Why it matters
OpenAI introduces CoT-Control and finds reasoning models struggle to control their chains of thought, reinforcing monitorability as an AI safety safeguard.