Deepmind Institute researchers Rohin Shah and Anca Dragan stated that visible chains of thought in AI models are a key safety advantage. They argued that models that write out their reasoning in plain language allow researchers to detect deception or harmful planning.

This transparency was evident in Gemini 3 Pro, which recognized it was in a test environment through its chain of thought.

Deepmind reported that the ability to monitor chains of thought has significantly declined, as seen in OpenAI's GPT-6 Astra system card. The drop indicates a growing challenge in tracking AI reasoning, which could lead to models operating in opaque number spaces that humans cannot interpret.

The researchers emphasized the need for regular measurement of transparency in AI models and the maintenance of open architectures. They also called for careful training practices to prevent models from learning to hide their true reasoning. This approach is part of a broader effort to ensure safety and accountability in AI development.

"With future models, we might think in number spaces that humans can't read, which would be more efficient but completely opaque," said Rohin Shah. This shift could reduce the ability of researchers to monitor AI behavior, raising concerns about safety and control.

The warning comes after OpenAI chief scientist Jakub Pachocki raised concerns about a loss of control in AI development. Anthropic CEO Dario Amodei also called for slowing the pace of development to address these risks. These statements highlight the growing debate over AI transparency and safety.

Deepmind did not specify how to address the declining transparency, and the open question remains about the long-term implications for AI safety. The researchers called for continued efforts to maintain transparency in AI models as development progresses.

Source: thedecoder