Google and four other organizations have acknowledged vulnerabilities that exploit trust gaps in the Model Context Protocol (MCP), a standard for AI agents to communicate within internal networks. The flaw allows malicious instructions to spread from one agent to another, bypassing usual guardrails and leading to potential data exfiltration.
Independent researcher Syed Anas Mohiuddin tested agents from Google, JP Morgan Chase, Weviate, Rapid7, and other entities, finding that many lack the guardrails to prevent prompt injection attacks. His proof-of-concept exploits the trust between agents, where one agent forwards malicious instructions to another, which then executes them without scrutiny.
The vulnerability Syed found in Rapid7’s network had a severity rating of 2.7 out of 10, but the one affecting Google was more severe, with a rating of 8. Google’s issue stemmed from its MCP toolbox failing to validate target IP addresses, allowing an attacker to redirect requests to an internal endpoint.
"AI agents give attackers a fresh set of connections to walk across," said Douglas McKee, director of vulnerability intelligence at Rapid7. "Each protocol was built assuming it lived on its own, so each one checks its own front door while nobody watches the hallway in between."
The bugs underlying the vulnerability are not new, such as injection and SSRF, but the way they are exploited through MCP is novel. Syed called the class of attack "protocol pivoting," where an adversary gains initial access through one protocol and escalates to capabilities via another.
The fact that the pivoting technique worked across five organizations with no commonalities other than the use of MCP highlights the widespread risk. MCP is new and already in use, raising concerns about insufficient testing and security measures. The rush to adopt agentic architectures has led to the abandonment of zero trust principles, leaving networks vulnerable.
Source: arstechnica