skip to content
The Weighted Average

Wire

Nvidia targets sub-300ms voice agents

Nvidia opened early access to Nemotron 3 VoiceChat, a 12-billion-parameter full-duplex speech-to-speech model that unifies automatic speech recognition, language reasoning, and text-to-speech in one architecture. Its early-access page describes open, inspectable weights for enterprise deployment, while Nvidia’s technical overview says the model targets sub-300ms end-to-end latency and processes 80ms audio chunks faster than real time. Voice builders should benchmark the single-model path against their cascaded stack, especially where interruption handling and orchestration failure matter; it extends the case for audio-first AI interfaces, but early access is not a production SLA.