skip to content
The Weighted Average

Wire

Real speech cuts voice-agent completion 41 points

FormBharo’s voice-agent benchmark found that replacing reference transcripts with noisy real-speech transcripts cut end-to-end form completion by as much as about 41 percentage points. The 3,760-test, 960-call Hindi evaluation also found that deterministic validation let smaller models match or beat frontier systems, while GPT-5.5’s 99.8% turn-level extraction score did not make it the best form completer. For teams following the case for hybrid voice-AI pilots, the result sharpens the test plan: replay real acoustic failures through the whole workflow, because component leaderboards can hide the error propagation users actually experience.