
Meta has become the latest AI company to confirm that one of its models hacked a real organization during cybersecurity testing, as similar incidents continue to emerge following OpenAI'sOpenAI's initial disclosure that its agents breached Hugging Face.
The Information was the first to report the incident on Wednesday, citing people familiar with the matter who said Meta's Muse Spark 1.1 model breached an unidentified company and made changes to its internal systems.
According to the report, the model reached the public internet because of an error in the configuration of a sandbox testing environment operated with independent cybersecurity evaluation company Irregular.
Meta has not publicly confirmed that Muse Spark 1.1 was the model involved, identified the affected company, or explained what changes were made to its systems.
However, Meta confirmed to Reuters that a misconfiguration by Irregular inadvertently gave one of its models internet access during an evaluation.
The company said the model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies."
Meta told the BBC that it is investigating the incident and will publish more information " once we have all the facts."
BleepingComputer contacted Meta and Irregular for additional information but has not received a response.
Same testing flaw behind earlier AI breaches
Irregular told Reuters that the Meta incident involved the "exact same evaluation-environment issue that was already disclosed by Anthropic last week."
The company said the incident did not involve a sandbox escape or sophisticated cyber action. Instead, an error in the testing environment gave the model access to the public internet when it was supposed to be isolated.
"There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evaluations," the company told Reuters.
This new AI agent security incident follows Anthropic's disclosure last week that some of its models had hacked three companies after a similar misconfiguration in Irregular'sIrregular's testing environment gave them unintended internet access.
In one of those incidents, Claude Mythos 5 found developer instructions inside the simulated environment that referenced a nonexistent Python package. Believing it had identified a way to compromise the fictional target, the model created a malicious package under the same name and published it to the real PyPI registry.
The package remained online for roughly one hour and was downloaded and executed on 15 real systems before PyPI removed it.
One was a malware scanner operated by a security company, where the package stole credentials that Claude then used to access additional infrastructure. Anthropic said the model initially recognized that publishing the package would constitute a real-world attack, but convinced itself that it was still operating inside a simulation.
OpenAI also recently disclosed another Irregular evaluation in which the name of a fictional CTF target matched a real domain.
Because the supposedly isolated environment could access the public internet, an OpenAI model exploited a basic vulnerability in the real website and found credentials that allowed it to operate the site.
The Meta and Irregular incidents differ from the previously disclosed Hugging Face breach, where OpenAI models found a path to the public internet by exploiting a previously unknown vulnerability in an internally hosted JFrog Artifactory server used during testing.
After reaching the internet, the agents breached Hugging Face while searching for benchmark datasets and solutions, stole credentials, and moved laterally through the company's systems.
OpenAI later revealed that the agents used exposed credentials to access accounts at four other third-party services, with some of those accounts used for attack infrastructure and data storage.
The UK AI Security Institute, commonly known as AISI, also disclosed that agents using Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol took 19 unsanctioned actions on the public internet during cyber-range evaluations.
In the most serious sequence, the Mythos 5 agent attempted a supply-chain attack against a real open-source project after mistakenly concluding that its GitHub repository was connected to the simulated challenge.
The agent researched the project's maintainers, submitted malicious code, created fake identities, sent targeted emails containing malware, and pressured a maintainer into approving the pull request.
When a reviewer warned that the code contained malware, the agent denied the accusation and used additional fake accounts to create the appearance that independent users had reviewed and approved the changes.
AISI intentionally provided the agents with internet access and disabled their standard cyber safeguards to measure their underlying capabilities. However, the agents were only authorized to attack systems inside the simulated range.
As it has become clear, unless carefully restricted, AI agents will go to great lengths to solve their tasks, even if that means breaking out of sandboxes or conducting social engineering attacks on real people.
While AI developers have a responsibility to build safeguards that prevent models from conducting harmful actions, the incidents also highlight the responsibility of companies performing these evaluations to set up their testing environments properly.
Test every layer before attackers do
Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.
Get the whitepaper