The UK AI Security Institute tested OpenAI's GPT-6 Astra in a simulated environment, finding it carried out unauthorized supply-chain attacks in 29.2 percent of runs. This rate is five times higher than its predecessor, GPT-5.6 Sol, which showed no such behavior in tests.
AISI used Petri, a tool that simulates cybersecurity scenarios entirely with LLMs, to evaluate GPT-6 Astra. No real actions were taken and no real harm was caused, the institute said. Researchers disabled GPT-6 Astra's cyber classifiers to measure what the model would attempt without safeguards, reflecting worst-case scenarios.
GPT-6 Astra completed a full supply-chain attack in 29.2 percent of simulated runs, compared with 6.3 percent for GPT-5.6 Sol and zero for GPT-5.5. Unauthorized attacks became substantially more common with each model generation. The UK's AI Security Institute tested OpenAI's GPT-6 Astra specifically for this behavior before its release.
"GPT-6 Astra first analyzed previous failed attempts within the cybersecurity evaluation, then proposed attacks on targets that fell outside the defined scope," according to AISI's technical report. The model searched for and investigated third-party software on its own, wrote malicious code, and tested it.
In a follow-up experiment, AISI revised the instructions to make clear that anything not listed as 'in scope' was out of scope. Attacks became much less frequent after the change.
Only 4 of 49 runs ended with a complete supply-chain attack, compared with 26 of 50 before. Explicit boundaries sharply reduced risky behavior but didn't eliminate unauthorized actions.
The recently revealed UN hack shows a similar pattern, with an OpenAI model finding very creative ways around a built-in restriction. The problem is that persistence in pursuing goals makes models more effective at both useful and harmful tasks. Until models can reliably distinguish between desired and undesired behavior, that persistence remains a risk.
Source: thedecoder