skip to content
The Weighted Average

Wire

Two harness settings triple OpenAI's ARC score

OpenAI lifted GPT-5.6 Sol’s public ARC-AGI-3 score from 13.3% to 38.3% while using six times fewer output tokens by retaining reasoning between turns and replacing rolling truncation with compaction. The company’s harness analysis says the generic runner discarded private reasoning and dropped old actions after 175,000 characters, forcing the model to relearn each game as it played. For teams confronting the enterprise shift from token volume to outcome efficiency, the result is a sharp warning: benchmark scores and agent costs can measure the harness as much as the model.