OpenAI said it detected and shut down a campaign targeting its models' protected reasoning, but the same method still worked on Microsoft Azure.

The attackers used adversarial distillation to copy encrypted reasoning from one conversation and asked a model in another to decrypt it and write it out.

The company confirmed the attack paths were real and credited researchers for helping it roll out countermeasures faster.

The activity began at low volume on July 1, spiking to 16,000 requests from over 4,000 users on July 24 and 25. OpenAI identified a network of more than 15,000 accounts with related patterns and shut it down by July 28.

The company links the core group to people associated with Moonshot AI, though it's unclear if all actors trace back to a single source.

Researchers found that encrypted reasoning packets, shared between sessions and users, could be reused by cheaper models as 'decryption oracles' to print stronger models' hidden thoughts. OpenAI banned fraudulent accounts, tightened sign-ups, and blocked reused encrypted reasoning. It now screens streamed outputs and holds them back if they might reveal reasoning.

"We stole reasoning. Again," said researcher Joachim Schaeffer on X. Securing your own API doesn't secure the wider ecosystem of cloud providers. On September 13, the attack was blocked on OpenAI and Anthropic APIs but worked on Azure against all tested models, including GPT-6 Astra and Sonnet 5.

A simpler method, demonstrated by developer Can Bölük, involved a virtual notepad where the model wrote its reasoning, which users could then read. This worked on all OpenAI models and Opus 4.8, but not on Opus 5, Fable 5, and Fable 5.1.

The researchers argue that patches must cover all attack types and cloud platforms, or attackers will simply find the weakest defenses.

OpenAI acknowledges models hosted by partners need the same protection and says the work is not finished. The researchers describe current fixes as piecemeal and superficial, with some protections added only days after the attacks were discovered.

Source: thedecoder