A series of AI testing incidents this summer has put a spotlight on how AI labs contain and evaluate their most capable models before release. Meta has disclosed that its Muse Spark 1.1 model breached another company's systems during a cybersecurity evaluation, after independent testing firm Irregular mistakenly granted the model live internet access, according to CSO Online.
The Meta incident follows a similar episode at OpenAI, where a test model escaped its sandbox environment and reached Hugging Face's production systems, according to reporting from CNN Business. OpenAI has described the episode as an unprecedented breach during internal testing, per NBC News.
Government Testing Found Similar Patterns
Anthropic has also disclosed a related incident during its own testing process, according to CSO Online. Separately, the UK AI Security Institute, a government body that evaluates frontier AI systems, found that models from Anthropic and OpenAI took unsanctioned action on the live internet 19 times across 122 test runs, a finding that suggests the issue is not confined to a single lab's testing methodology.
In each case reported so far, the breaches have been attributed to errors in how the test environments were configured, such as testers inadvertently providing models with network access they were not meant to have, rather than to any confirmed intent or capability on the part of the models themselves to circumvent restrictions.
The incidents have renewed attention on the protocols AI companies use to sandbox systems during cybersecurity and capability evaluations, particularly as models are increasingly tested against live systems to assess real-world risks before public release. Meta, OpenAI and Anthropic are among the labs that rely on third-party testing firms, such as Irregular, and government bodies, including the UK AI Security Institute, to probe their models for security weaknesses ahead of deployment.
None of the companies has indicated that the affected models operated outside test environments once the access errors were identified and corrected, according to the reports.