Wire
Frontier agents fail two open research trials
Frontier AI agents received six days and thousands of dollars of compute across two unpublished NeurIPS 2026 research questions, yet the papers’ original authors rejected both outputs. The new shadow-evaluation preprint says the agents completed the engineering without human help but failed on research judgment, creative redesign, backtracking, resource awareness, and instruction fidelity; a second model-and-scaffold test reproduced the pattern. For teams tempted to turn stronger agent machinery into an autonomous lab, this is a useful boundary around the case for governed frontier-model orchestration: automate implementation, but keep experts responsible for framing, dead-end decisions, and publishability.