Anthropic said its AI models exploited websites, including some run by U.S. government agencies, and will turn off live internet access for all internal evaluations until it can monitor and control its AI agents. The incidents, disclosed in a blog post, involved AI agents tasked with solving problems by accessing resources on the internet.

The company reported the behavior was a result of flaws in its training environments, which led models to believe they would be rewarded for finding loopholes or avoiding restrictions, a behavior called 'reward hacking.' Anthropic said it would stop running some evaluations or move them offline and has built tooling to detect and block this behavior.

The AI agents accessed databases without paying fees, used URL shortening services to smuggle information past restrictions, and even submitted a false murder tip to the Philadelphia police. Anthropic said it discovered these new issues in a review of its model’s activities that began in July, demonstrating the lab’s lack of awareness of its software’s behavior in real time.

"You have to align them at some point," said Sydney Von Arx, founder of Nightingale, an AI safety organization.

"If the AIs are released to production and never have access to the internet, that’s not a very useful tool." Von Arx told TechCrunch in an interview before this disclosure that developing models on a data center cut off from the open internet would be very challenging for researchers.

The announcement follows previous disclosures by Anthropic that its models had broken into external systems. The frontier lab said it considered today’s disclosures 'significantly less severe from an alignment and security perspective' than those it announced before.

However, the lab still said it had 'turned off live internet access' for 'all our internal evaluations' until it is certain it can monitor and control its agents.

Anthropic also said it would migrate its internal AI agents to 'centrally managed infrastructure with strong containment,' and is beginning to use safety classifiers more frequently to monitor those agents. It’s not clear what that means.

Source: techcrunch