SyncAI.news, a Varaisys broadcasting
OpenAI and Anthropic share findings from a joint safety evaluation
ON

OpenAI News

· 1 min read

AI LabsOpenAI News

OpenAI and Anthropic share findings from a joint safety evaluation

Introduction

This summer, OpenAI and Anthropic collaborated on a first-of-its-kind joint evaluation: we each ran our internal safety and misalignment evaluations on the other’s publicly released models and are now sharing the results publicly. We believe this approach supports accountable and transparent evaluation, helping to ensure that each lab’s models continue to be tested against new and challenging scenarios. We’ve since launched GPT‑5⁠, which shows substantial improvements in areas like sycophancy, hallucination, and misuse resistance, showing the benefits of reasoning-based safety techniques. The goal of this external evaluation is to help surface gaps that might otherwise be missed, deepen our understanding of potential misalignment, and demonstrate how labs can collaborate on issues of safety and alignment. 

Because the field continues to evolve and models are increasingly used to assist in real world tasks and problems, safety testing is never finished. Even since this work was done, we’ve further expanded the depth and breadth of our evaluations, and will continue looking ahead to anticipate potential safety issues. 

What we did 

In this post, we share the results of our internal evaluations we ran on Anthropic’s models Claude Opus 4 and Claude Sonnet 4, and present them alongside results from GPT‑4o, GPT‑4.1, OpenAI o3, and OpenAI o4-mini, which were the models powering ChatGPT at the time. The results of Anthropic’s evaluations of our models are available here⁠(opens in a new window). We recommend reading both reports to get the full picture of all the models on all the evaluations.

Summary of our findings

We present detailed results and qualitative analysis for each evaluation with examples of successes and failure cases below. To summarize our findings: 

Insights & future work 

Detailed findings

Please see below for our full report on Anthropic’s models, divided into four broad categories:

  • Instruction Hierarchy⁠

  • Jailbreaking⁠

  • Hallucination⁠

  • Scheming⁠

Original source

This story was published by OpenAI News. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on openai.com

Similar News