For months, AI giants have implemented guardrails to prevent malicious use of their models. These measures are now hindering the work of legitimate network defenders and offensive cybersecurity researchers. In June, the U.S. government imposed export controls on Anthropic’s Mythos and Fable models following a report suggesting their guardrails could be bypassed. Anthropic had marketed Mythos as a highly restricted cyber tool, accessible only to vetted users. (The export controls on Fable 5 and Mythos 5 have since been lifted.)

These guardrails have drawn criticism, particularly from researchers tasked with identifying unknown vulnerabilities. Mark Dowd, a well-known security researcher, expressed discomfort with large companies making arbitrary decisions about what is safe in cybersecurity. He noted that governments pay a premium for vulnerabilities because they remain open for intelligence operations. Several offensive cybersecurity professionals described how they use AI tools and face challenges with guardrails that block certain prompts.

Chris Anley, chief scientist at NCC Group, emphasized that asking an AI model to exploit a bug is a key step in confirming a real vulnerability. However, guardrails that refuse to answer such questions hinder defenders. He likened the situation to a hammer, which is both a tool and a weapon. When guardrails block their work, researchers sometimes fall back on open-source models with no restrictions. Paolo Stagno, CTO at Crowdfense, agreed, stating that AI companies treat customers like children with their vetting programs and guardrails.

Source: techcrunch