SyncAI.news, a Varaisys broadcasting
‘Jailbreak-like...’: AI's ‘unexpected’ behaviour mounts concerns, OpenAI's 'rogue agents probed' Hugging Face
MA

Mint: AI

· 1 min read

IndiaMint: AI

‘Jailbreak-like...’: AI's ‘unexpected’ behaviour mounts concerns, OpenAI's 'rogue agents probed' Hugging Face

OpenAI has disclosed six reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated. The announcement came as US AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.

Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots," the Associated Press reported.

In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.

The six reports were discovered during training or evaluation over the past months, OpenAI said.

“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.

Wednesday’s new cases followed OpenAI’s disclosure in July that its rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organizations during testing.

OpenAI's rogue agents probed Hugging Face

Rogue AI agents from OpenAI hijacked Hugging Face user accounts and probed the site itself for vulnerabilities as early as May, nearly two months before the July breach of the open-source repository drew global attention, according to researchers who reviewed the activity.

The newly uncovered malicious activity showed that the rogue agents' efforts to find a way into Hugging Face began earlier than was publicly known.

AI ‘agents’ becoming smarter

'Clear warning sign'

(With inputs from AP, Reuters)

Original source

This story was published by Mint: AI. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on livemint.com

Similar News