Wire
Databricks harness lifts grounded reasoning 24 points
Databricks Genie improved matched frontier models by an average 24 percentage points on the new 90-question OfficeQA Pro V2 benchmark. The vendor-run test spans roughly 1,400 Treasury PDFs and 120,000 pages; Claude Fable 5 gained 14.4 points in Genie while its rollout cost fell about ninefold versus Claude Code, excluding Genie’s one-time corpus-parsing cost. The result reinforces the case for benchmarking the whole model-harness system: for document agents, parsing and loop control may buy more than switching frontier models.