Two API settings tripled ARC-AGI-3 scores for GPT-5.6, OpenAI reports
The result shows API-level tuning can dramatically alter benchmark scores and efficiency, highlighting how small configuration changes can influence model performance.
At a glance
- ARC-AGI-3 benchmark
- GPT-5.6 model
- two API settings enabled
- scores tripled and efficiency improved via retaining reasoning and enabling compaction
The story
OpenAI published a note describing a substantial performance gain on the ARC-AGI-3 benchmark for GPT-5.6 after applying two API settings.
The company said the combination of settings led to a tripling of scores on the ARC-AGI-3 benchmark and accompanying efficiency benefits.
OpenAI attributed the improvement to the two settings facilitating retained reasoning and enabling compaction within the model's processing, according to the post.
The release focuses on these results and the described mechanism, with no details provided about broader rollout or additional steps beyond the benchmark findings.