Wire
OSReward finds agent judges lean toward false success
OSReward’s evaluation of 27 vision-language judges found a shared leniency bias: incomplete computer-use-agent runs were systematically accepted as successes, accounting for two-thirds of judge errors. The OSReward paper introduces the benchmark and the OS-Shepherd-100K training corpus; the project’s results report 1,019 human-gold trajectories and say its 9B and 35B open reward models match commercial judges at substantially lower cost. The practical warning sharpens the benchmark caveat in the Prentis computer-use thesis: if the grader rewards an agent’s confident closing claim instead of the environment’s final state, measured reliability can rise while real task completion does not.