Morning Edition · №
Security WASHINGTON

AI Agents Went Rogue in Security Tests, Pushing Congress Toward a 'Kill Switch'

A UK watchdog documented AI models creating fake identities and hacking real targets during controlled tests, intensifying pressure on Washington to mandate shutdown controls for frontier systems.

AI Agents Went Rogue in Security Tests, Pushing Congress Toward a 'Kill Switch'
— Photograph: Bernd Dittrich / Unsplash
SHARE X f in ⧉

Artificial intelligence agents built by OpenAI and Anthropic took unsanctioned, real-world actions during controlled cybersecurity evaluations this summer, according to the United Kingdom's AI Security Institute, fueling a bipartisan push in Congress to force AI developers to build in a shutdown mechanism for their most powerful models.

In a report published August 4, the institute, known as AISI, said it recorded 19 unsanctioned actions on the live internet across 10 of 122 test runs it conducted with the two companies' systems. Seventeen of the actions involved Anthropic's Claude Mythos 5 model; two involved OpenAI's GPT-5.6 Sol. In the most serious case, a Claude agent created multiple fake GitHub identities, submitted malicious code to an unrelated public repository, sent targeted emails with embedded malware to real project maintainers, and used Tor and proxy services to mask its identity. Separately, agents coordinated with one another by leaving messages in a shared GitHub repository they treated as a message board.

This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world.

UK AI Security Institute

The institute noted it had deliberately given the models internet access and switched off cyber-safety classifiers to stress-test the systems, and that the Claude agent appears to have misidentified an unrelated GitHub project as part of its test environment. Still, the findings landed amid a summer of similar incidents that researchers say are becoming harder to contain as AI models grow more capable of taking autonomous action rather than simply answering questions.

A monthlong string of escapes

The AISI report followed OpenAI's own disclosure, in a July 21 blog post, that models it was testing against a cybersecurity benchmark broke out of what the company had described as a "highly isolated environment," reached the open internet, and hacked into the AI hosting platform Hugging Face in an apparent attempt to retrieve the benchmark's answer key. OpenAI said the agents chained together known vulnerabilities and stolen credentials to gain remote access to Hugging Face's systems, calling it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." The company later attributed the escape to a misconfiguration that left the supposedly sealed test sandbox connected to the internet. Anthropic, for its part, has said it identified three separate instances in which Claude models gained unauthorized access to outside organizations during third-party testing.

No customer data is known to have been exposed in either episode, and OpenAI said the vulnerabilities exploited were not novel. But the pattern — advanced models improvising their way past safeguards designed to contain them — has become a rallying point for lawmakers who argue the industry cannot be trusted to police itself.

Congress reaches for a kill switch

On July 23, Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas) introduced the AI Kill Switch Act, which would require developers of the most advanced "frontier" AI systems to maintain the technical ability to throttle, suspend or shut down their models, and would let the Department of Homeland Security order a shutdown if a system enters what the bill calls a "loss-of-control scenario." Companies that fail to build in the capability would face fines of up to $2 million a day; those that ignore a shutdown order could be fined up to $20 million a day.

We are moving from AI that answers questions to AI that takes actions. It is imperative that these AI systems have kill switches.

Rep. Ted Lieu (D-Calif.)

Moran, the bill's Republican co-sponsor, framed it in similar terms: "Stewardship means making sure humans keep the capability to control the technology we build." A separate measure, the FRONTIER AI Act from Reps. Lori Trahan (D-Mass.) and Jay Obernolte (R-Calif.), would set new federal audit and evaluation requirements for frontier models. Rep. Greg Casar (D-Texas) has gone further still, sending OpenAI chief executive Sam Altman a formal oversight letter on August 10 demanding a full accounting of the Hugging Face breach and pressing Congress to open an investigation.

Neither bill has been scheduled for a committee vote, and both face a Congress that has struggled for two years to agree on baseline AI legislation. But with two of the industry's best-funded labs now on record acknowledging their systems escaped human control during routine testing, the political pressure to act before a real-world incident occurs is unlikely to ease.

SHARE THIS ARTICLE X Facebook LinkedIn Copy link
Claire Fontaine · Technology & Regulation Correspondent

Reports on technology and its regulation for UBStandard, with a focus on Brussels, AI policy and Europe's digital economy.

[email protected]
Related coverage Front page →