skip to content
The Weighted Average

Wire

A deadline cuts an AutoML win rate to 34.3%

A nominally 60-second AutoML system consumed a median 120 seconds, and its apparent win rate fell from 59.4% to 34.3% when researchers enforced the deadline and stopped selecting candidates on the test split. The Winning by Peeking audit found no significant pairwise advantage after equalizing compute and moving selection to validation data. The lesson sharpens the procurement case for pinned harnesses and raw traces: externally enforce wall-clock budgets and isolate the test set, or a precise leaderboard can still rank protocol defects.