skip to content
The Weighted Average

Wire

Samaya opens a 220-query finance-agent benchmark

Samaya released FrontierFinance, an open benchmark with 220 investment-workflow queries and 11,543 expert-written rubrics, and reports its agent qualifying 50.8% of rubrics versus Claude Fable 5’s 49.2% at roughly one-quarter the inference cost; the dataset and grading code are public. The same model performs better behind a specialized finance harness than web search alone, extending the argument that agent harnesses—not just base models—set practical capability. Finance teams now have a reproducible starting point for vendor trials, but should rerun the grader on their own filings, data permissions, latency, and error costs before treating Samaya’s self-reported lead as a buying verdict.