OpenAI recently identified and disrupted a coordinated campaign aimed at extracting protected reasoning from its models. The earliest observed activity occurred in the first week of July, with high-volume spikes on July 24 and 25 consisting of 16,000 requests using a relevant extraction pattern from over 4,000 users.
The campaign involved operators attempting to extract protected reasoning in novel ways, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content. Independent security researchers also brought related cross-model and conversation-compaction vulnerabilities to OpenAI’s attention through responsible disclosure.
OpenAI confirmed the attack paths identified by the researchers were real and used their findings to understand the broader attack class and accelerate mitigations. The activity began on July 1, initially at a low volume, but escalated to a cluster of more than 15,000 users by July 28, which was fully disrupted by the company.
"We mitigated this recent distillation campaign through a combination of account enforcement, technical controls, and partner coordination," said OpenAI. The company banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks.
The disruption of the campaign highlights the broader security challenge of adversarial distillation, which poses safety and national security risks. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs. OpenAI emphasized the need for coordinated industry efforts to address this shared security challenge.
OpenAI did not specify whether all operators observed originated from a single actor, and noted the evolving nature of adversarial distillation as frontier models improve. The company continues to improve tool defenses, classifier coverage, model refusals, and propagate relevant controls across cloud partners.
Source: openai