Wire
ORCA-bench caps on-call agents at 25.3%
ORCA-bench gave five frontier agents 1,079 production-style root-cause tasks, and the best managed 25.3% accuracy on medium difficulty and 10% on hard cases, according to the benchmark’s July 30 paper. The public testbed contains six days and 50 GB of metrics, logs, traces, and source code; the weakest model hallucinated an implausible cause in 40% of incident reports, while removing source access hurt every metric. That result is a sharp boundary on managed-agent platform promises: builders can automate evidence collection now, but production remediation still needs a human incident commander and explicit approval gates.