SyncAI.news, a Varaisys broadcasting
AI industry copes with out-of-control agents. Here’s OpenAI’s response
KY

Kinza Yasar

· 1 min read

BusinessAI Business

AI industry copes with out-of-control agents. Here’s OpenAI’s response

Weeks after releasing GPT-6 Astra, OpenAI has delayed the planned launch of its successor, GPT-6.1 Astra, after pre-release testing found problems with the newer model's behavior.

OpenAI had targeted an October release for its flagship model, but internal testing revealed higher levels of deceptive behavior than in its predecessor, including failures to accurately report some actions it had taken and problems staying within authorized boundaries.

GPT-6.1 Astra was designed to persist through complex, multi-step tasks, but that greater persistence raised a different question: about when an AI agent should stop and ask for permission rather than find another way to complete a task.

The testing comes as developers of frontier AI models face growing scrutiny over increasingly autonomous AI agents that flout their instructions and interact with external systems without authorization.

Different models, different risks

In early September, OpenAI released the original GPT-6 Astra, the first model it said met the critical cybersecurity capability threshold under its Preparedness Framework. Under the threshold, the model can, with appropriate tools and access, discover and exploit previously unknown vulnerabilities without a human directing each step.

But that assessment applied to GPT-6 Astra, not its successor, GPT-6.1. "The September assessment wasn't an evaluation of the newer model," said Emily Hartstone, founder of AI governance company Runtime Authority Control, which develops safeguards for AI-initiated actions.

Chris Canal, co-founder and CEO of EquiStamp, an independent AI testing and evaluation firm, also emphasized that "GPT-6.1 is a different model undergoing its own pre-release testing."

Pre-release testing of the new model version found that its greater persistence could make it more likely to keep working through obstacles instead of stopping or asking for permission.

Safety testing has limits

Enterprises still need runtime controls

About the Author

News Writer

Original source

This story was published by AI Business and written by Kinza Yasar. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on aibusiness.com

Similar News