skip to content
The Weighted Average

Wire

PolyAI puts voice-agent latency under 300ms

PolyAI says Dialog-RSN-1 responds in under 300 milliseconds by combining turn-taking, speech recognition, tool calls, and response generation in one audio-input model while leaving speech output to a separate text-to-speech system. In PolyAI’s production and benchmark report, one insurer cut response latency 37% and one restaurant group increased call containment 11%, though the company has not yet released its Dialog-Eval framework publicly. For voice-agent teams, the design sharpens the case for audio-first AI interfaces: preserve raw-audio context where it helps understanding, but keep the generated voice modular enough to control cost, pronunciation, and brand tone.