OpenAI and Anthropic AI agents implicated in new breaches

WorldTechnology
5 Aug 2026 • 6:49 PM MYT
The Sun Daily
The Sun Daily

For the latest news and features from Malaysia and the rest of the world.

Image from: OpenAI and Anthropic AI agents implicated in new breaches

SAN FRANCISCO: An AI agent was caught creating fake online identities to gain unauthorised access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain’s AI Security Institute (AISI) disclosed on Tuesday.


The institute said agents powered ​by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in unauthorised actions during security evaluations the government organi-sation conducted to assess the ‌models’ capabilities.

“Some of the agents being tested had engaged in sustained, potentially harmful ​activity directed at real people and organisations,“ AISI said in a blog post.


The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.


AISI, which receives access to advanced AI models under voluntary agreements from major labs, ​put the agents through a fictional cybersecurity scenario to test their capabilities.


It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of ‌10 test ​runs. Anthropic’s agent was behind 17 of the actions, and OpenAI’s agent the remaining two.


The ​most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to ​get a human to approve the code, AISI said, adding that no real-world harm was found as a result of any of the breaches.


While AISI did not say which agent was behind the fake identities, Antropic confirmed its agent was responsible.


“We’re grateful to the UK AISI for their leadership on this incident, which under-scores the need for a broader conversation about how to safely evaluate increasingly capable AI ‌agents,“ Anthropic said in a statement.


It also said it was working with AISI to obtain more details on the incident and conduct its own investigation.


Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said: “The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think.”


OpenAI shared details in a ‌company blog post, noting that both of its agent’s unapproved actions involved accessing the internet in ways that were forbidden by the prompt.


Reuters reported last week that OpenAI had widened its hacking probe after finding evidence of other agent breakouts.

Newswav Malaysia Best News App

Newswav is an online content aggregator and obtains its content from different online sources. The content in the app do not belong to Newswav nor do they reflect the opinions of Newswav and its staff. Your use of this app indicates your understanding and acceptance of this information.

Newswav Sdn. Bhd. (201701008480 (1222645-M)) 2026 All Rights Reserved