skip to content
The Weighted Average

Wire

GPT-5.6 Sol scores 53.6 on Agents' Last Exam

OpenAI launched the three-model GPT-5.6 family, with Sol scoring 53.6 on Agents’ Last Exam, 13.1 points above Claude Fable 5, and introduced an ultra setting that coordinates four agents by default. The release turns parallel-agent compute into a first-party capability lever, sharpening the price-performance contest framed by Claude Opus 5’s frontier-coding economics. Builders should benchmark total successful-task cost rather than token price alone, because the premium mode deliberately spends more tokens to trade parallel compute for stronger results and lower latency.