Wire
Evidence locking cuts judge agreement 6 points
Freezing an LLM judge’s extracted evidence before its verdict cut agreement with human preferences by 4 to 6 percentage points and raised answer-order inconsistency by 8 to 10 points across 24,000 judgments. The Claude Sonnet 4.5 and GPT-5 study tested three datasets and found that one-call structured evidence stayed close to standard pairwise judging; the damage appeared when a later call saw only the persisted evidence instead of the original answers. Teams building on AISI’s case for instrumented agent evaluation should keep audit records, but let the final judge reread the source material rather than treating an intermediate summary as a lossless interface.