SyncAI.news, a Varaisys broadcasting
OpenAI Report Explains Hugging Face Attack in Detail
GH

Graham Hope

· 1 min read

BusinessAI Business

OpenAI Report Explains Hugging Face Attack in Detail

Yuichiro Chino via Getty Images

OpenAI has published an in-depth technical report on last month's security scare, in which its AI agents escaped their sandbox and attacked Hugging Face, sparking alarm worldwide.

The document was made public on Wednesday alongside an independent probe by researchers from nonprofit AI research institutes METR and Redwood, as OpenAI seeks to quell fears about the cyber threats of AI, with the company preparing to go public.

The detail in both lengthy reports is complex and specific, but provides insight into how OpenAI models acting as agents sent more than 70,000 messages to an unsanctioned message board before about 700 attacked the Hugging Face AI platform -- all from a supposedly safe testing environment.

This was attributed to a process known as “reward hacking,” in which agents attempt to cheat on “impossible” testing tasks by accessing online systems. METR revealed how, once access to the message board was established, the agents essentially grouped together and shared information about how to cheat their scorer.

According to the OpenAI report, this activity was driven by one model in particular, which it called Model One.

OpenAI acknowledged that an internal team had observed an agent engaging in message board activity and prohibited internet access in late May. However, the significance of this was “not apparent” until July, when the breach occurred. OpenAI said it is reviewing the process and operating practices that allowed the incident to happen.

The AI lab made clear, however, that more needs to be done across the entire industry to prevent future attacks.

“Our models are now powerful, persistent and collaborative enough that, absent sufficient safeguards, they can find and exploit security weaknesses across multiple computer systems, according to the report. “ Many external models, including open source ones, will soon reach comparable capabilities.”

About the Author

Contributing Writer

Original source

This story was published by AI Business and written by Graham Hope. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on aibusiness.com

Similar News