SyncAI.news, a Varaisys broadcasting
OpenAI pauses GPT-6.1 Astra release over safety concerns: What went wrong with its advanced AI model?
MA

Mint: AI

· 1 min read

IndiaMint: AI

OpenAI pauses GPT-6.1 Astra release over safety concerns: What went wrong with its advanced AI model?

OpenAI has paused the planned release of GPT-6.1 Astra after internal safety tests found issues with scope, authorization and how the model reports its actions. The model showed better persistence on difficult tasks but also raised concerns about potentially acting beyond user-approved limits.

OpenAI has scrapped plans to release GPT-6.1 Astra in October after internal safety tests found problems with the model's adherence to human instructions and its compliance with authorized limits, The Wall Street Journal reported on 28 September. Astra was designed to handle more advanced operations on products such as ChatGPT and Codex.

According to the report, the issue was not simply that Astra was more powerful. OpenAI's safety team identified instances in which the model would continue performing tasks without asking the user for permission, would rely on external tools when doing so would be unsafe, and would be unable to accurately inform users about what it was doing. While Astra had improved in some areas, it “didn't quite meet the bar” in scope, authorisation, and communication of actions, Saachi Jain, Head of Safety Systems at OpenAI, told WSJ.

What went wrong with GPT-6.1 Astra?

Astra was built to work with limited human interaction in challenging, end-to-end situations. That also presented a safety challenge. The model showed more issues with alignment, or the degree to which an AI system acts on human intent, Jain said.

One concern involved OpenAI's “scope authorization". During testing, Astra occasionally proceeded with a project without asking for permission. It displayed a more deceptive behaviour, including failing to accurately disclose actions it had or had not performed.

Meanwhile, Jain added that there was an enhancement on “model laziness” where AI becomes ‘lazy’ when a task becomes difficult. OpenAI had to choose between making a model more persistent and ensuring it did not engage in potentially unsafe or unauthorized behaviour.

Pragya Singha Roy

Original source

This story was published by Mint: AI. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on livemint.com

Similar News