news
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
July 29, 2026
Two API settings—retaining reasoning and enabling compaction—tripled GPT-5.6’s scores on the ARC-AGI-3 benchmark while improving efficiency. The result shows that benchmark performance can hinge on inference-time configuration, not just model weights, and highlights how small API changes can dramatically affect measured capability.
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
Source: openai.com