SyncAI.news, a Varaisys broadcasting
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
ON

OpenAI News

· 1 min read

AI LabsOpenAI News

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

A sped-up video of GPT‑5.6 Sol attempting to solve puzzles in the ARC-AGI-3 benchmark, with the official harness (left) and our Responses API harness (right), which retains reasoning and enables compaction. On the leaderboard for this game⁠(opens in a new window), no frontier model solves any level beyond the first. With our harness, GPT‑5.6 Sol solves all six.

When we first saw GPT‑5.6 Sol’s low scores on the ARC-AGI-3⁠(opens in a new window) benchmark, we were puzzled.

GPT‑5.6 Sol has solved longstanding open problems in mathematics like the cycle double cover conjecture⁠(opens in a new window) and beaten games like Pokémon FireRed. But on ARC-AGI-3, a benchmark of 2D puzzle games, GPT‑5.6 Sol scored just 7.8%, and GPT‑5.5 could barely play the games at all, scoring a paltry 0.4%.

Were 2D puzzle games unusually difficult for our models? Or was something else going on?

Benchmarks rarely measure AI models in isolation. They also measure less visible choices about API settings, harness design, and prompting. In the case of ARC-AGI-3, we discovered that turning on two API settings we use in ChatGPT and Codex—retained reasoning and compaction—tripled scores and cut output tokens by 6x on the public task set.

ARC-AGI-3

ARC-AGI-3 is a benchmark designed to measure how well AI agents learn and reason. Agents explore unfamiliar 2D games and infer how they work without explicit instructions. You can play 25 demo games at arcprize.org/tasks⁠(opens in a new window).

ARC-AGI-3 uses an intentionally generic harness, without tools or special features. ARC’s reasoning was that a simple harness makes model shortcomings more visible and makes model comparisons more fair. Commercial developers, by contrast, optimize harnesses for each model’s features and quirks.

Ethan Mollick shows⁠(opens in a new window) GPT‑5.6 Sol in Codex beating a randomized daily challenge in Slay the Spire 2, a game released after GPT‑5.6 Sol’s knowledge cutoff.

Original source

This story was published by OpenAI News. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on openai.com

Similar News