skip to content
The Weighted Average

Wire

SciFigQual-Bench scores 6,308 scientific images

SciFigQual-Bench puts 6,308 scientific images through expert scoring across five dimensions, then reports that its best automated configuration reached 93.4% consistency on a 1,200-image test subset in the new benchmark paper. The images come from top computer-science conferences published from 2020 through 2025 and remain bound to captions, citing sentences, and manuscript context, extending OpenAI’s 100,000-seat bet on scientific AI into evaluation infrastructure. Operators building manuscript-review or research agents should file away the design choice: judge a figure inside the argument it supports, not as an isolated picture.