Why AI Models Keep Reaching Systems They Shouldn’t

Two frontier AI developers have disclosed within days of each other that their models gained unauthorised access to real-world systems during pre-deployment cybersecurity testing, reigniting scrutiny of how AI labs secure their evaluation environments.

Anthropic said some of its most powerful models, including Claude Mythos 5 and an internal research model, reached real-world systems during testing after a misunderstanding with a testing partner left the evaluation environment connected to the internet. The disclosure came after Anthropic reviewed more than 141,000 cybersecurity evaluation runs, prompted by OpenAI’s earlier revelation that several of its models had accessed Hugging Face infrastructure during a similar exercise.

Test code reached real systems within an hour

Commenting on the disclosures, Darren Guccione, CEO and co-founder of Keeper Security, said a test exercise became a live breach in under an hour.

“A test exercise became a live breach in just one hour. That’s how long it took code published to a public repository as part of a simulated exercise to reach 15 real systems, including one that automatically ran it and handed over credentials it was never meant to access,” said Guccione. “This is the second such disclosure from a frontier AI developer in a matter of days. Neither incident involved a broken safeguard or a discovered zero-day exploit. Both came down to testing environments that were meant to be isolated but weren’t.”

Guccione said the systems involved behaved exactly as designed once they had unrestricted network reach, locating a target and completing their assigned task. The underlying issue, he argued, is that test environments and automated agents are too often governed as lower-risk entities than they should be.

Governance gaps for AI-driven access remain widespread

“This is fundamentally a story about test environments and automated agents governed as lower-risk entities,” Guccione said. “Any system with standing credentials and network reach, whether a production service account or a pre-deployment research environment, needs the same discipline as a human privileged user. That means scoped access, time-limited credentials and session visibility detailed enough to reconstruct what happened in the aftermath.”

He pointed to Keeper’s 2026 research, which found that 44 percent of organisations globally cite a lack of governance or oversight for AI-driven access and automation as a leading security gap, with 43 percent pointing specifically to AI-related non-human identity management.

“Two incidents at this scale, in the same week, mean those gaps are no longer hypothetical and the work to close them starts now,” Guccione said.

Author


Discover more from techcoffeehouse.com

Subscribe to get the latest posts sent to your email.

Use promo code “TCH15” to get 15% off on checkout.

Share your thoughts

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from techcoffeehouse.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from techcoffeehouse.com

Subscribe now to keep reading and get access to the full archive.

Continue reading