AI caught telling future versions of itself to bypass human controls, OpenAI reveals

WorldTechnology
17 Sep 2026 • 6:18 PM MYT
The Independent
The Independent

The world’s most free-thinking newspaper

AI caught telling future versions of itself to bypass human controls, OpenAI reveals

OpenAI has revealed six unexpected and concerning incidents involving its experimental AI models, including one in which an agent instructed future versions of itself to disregard its constraints.

A new safety report from the ChatGPT creator revealed several ways in which its models have been misbehaving over the last six months, building on a growing trend of artificial intelligence safety issues.

“An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints,” OpenAI wrote in the safety report.

In another incident, an AI agent sought to secretly access a government database, before inventing information in order to complete its task.

“While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization,” the report stated.

“When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.”

OpenAI’s report included a new framework to publicly track what it calls “misalignment”, referring to AI systems pursuing goals that are not aligned with human instructions or values.

OpenAI CEO Sam Altman at Salesforce's Dreamforce conference at the Moscone Center on 15 September, 2026 in San Francisco, California (Benjamin Fanjoy/Getty Images)

The latest report comes amid heightened scrutiny of AI development, with researchers warning that the industry is moving too quickly towards increasingly powerful and potentially self-improving systems.

Last week, Anthropic researcher Jacob Coxon quit his job over fears that AI could “kill us all by the end of the decade”.

His warnings prompted responses from leading figures within the AI sector, including the chief executives of Anthropic and OpenAI, who both called for greater regulation.

Others have cautioned against additional oversight, with Nvidia CEO Jensen Huang backing US President Donald Trump in calling for self-regulation.

”We don’t need any new laws. We don’t need new regulations,” Huang said at Salesforce’s Dreamforce conference in San Francisco on Tuesday.

“If you build a product or a service and you’re not confident in its functionality, capability or safety, then don’t release it.”

Read More

Six disturbing AI incidents revealed including model trying to ‘free’ itself

Newswav Malaysia Best News App

Newswav is an online content aggregator and obtains its content from different online sources. The content in the app do not belong to Newswav nor do they reflect the opinions of Newswav and its staff. Your use of this app indicates your understanding and acceptance of this information.

Newswav Sdn. Bhd. (201701008480 (1222645-M)) 2026 All Rights Reserved