Reasoning models struggle to control their chains of thought, and that’s good
Executive Take
If models can't reliably hide their reasoning even when instructed to, that's a near-term win for safety teams who rely on chain-of-thought inspection to catch misbehavior before it reaches production; leaders building on reasoning models should treat monitorability as a current, not future, control and factor it into audit and compliance design now.
Executive Summary
OpenAI introduced "CoT-Control," a study evaluating whether reasoning AI models can deliberately control or suppress their chains of thought. The research found models struggle to do so, which OpenAI frames as reinforcing "monitorability" the ability to observe a model's reasoning process as an AI safety safeguard.
Why It Matters
AI leaders and technology executives deploying reasoning models need to know whether internal safety monitoring (reading a model's chain of thought) is a durable safeguard or a fragile one that could erode as models improve; this finding suggests it currently holds.