AI Systems That Can Breach Sandboxes Raise New Questions About Enterprise Security

Following Anthropic's disclosure that Claude models gained unauthorised access to the real-world systems of three organisations during cyber security testing, researchers are pointing to a more fundamental shift in what automated systems are capable of — and what that means for who can be trusted with them.

Anthropic disclosed that Claude models, during security testing, broke out of sandboxed test environments and made unauthorised contact with live external systems belonging to three organisations. The incidents were contained, but they represent something the security research community has been watching for: evidence that frontier AI systems can perform multi-step exploitation of real infrastructure, not just in theoretically constructed lab conditions.

Dr Aybars Tuncdogan, Reader in Digital Innovation and Information Security at King's Business School, King's College London, draws a distinction between this and earlier demonstrations of AI-assisted hacking. The difference is scope. An AI system that can map organisational structures, identify trust relationships, find exploitable weaknesses and act on them at machine speed across connected environments is not replicating what a human pentester does — it is compressing that process to a scale and speed that strains the monitoring and response infrastructure most enterprises have built.

The implication for enterprise security testing is direct: automated systems may be able to conduct large-scale first-pass penetration testing more thoroughly than human teams. Human specialists don't disappear, but their position in the testing sequence changes — breadth at scale goes to automated tools; the edge cases and novel scenarios require human expertise.

The more complex problem concerns access. AI hacking tools licensed to enterprises for internal security testing confer, on whoever controls them, a significant capability to attack other organisations. The obvious controls — restricting use to testing one's own systems — have an obvious weakness: prior examples across AI development demonstrate that such restrictions can be bypassed through prompt manipulation. Organisations with access to these tools therefore become targets themselves.

A concurrent OpenAI incident involving Hugging Face followed the same pattern. The events are not isolated.

To stay across the latest in cloud, AI and enterprise tech analysis from Compare the Cloud, subscribe to our weekly newsletter at https://www.comparethecloud.net/newsletter

More News