gpt.buzz
Sign in

news

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

July 29, 2026

Two API settings—retaining reasoning and enabling compaction—tripled GPT-5.6’s scores on the ARC-AGI-3 benchmark while improving efficiency. The result shows that benchmark performance can hinge on inference-time configuration, not just model weights, and highlights how small API changes can dramatically affect measured capability.

How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.

Source: openai.com

← All news

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark · gpt.buzz