skip to content
The Weighted Average

Wire

Diffusion LLMs commit answers by 24% of decoding

LLaDA-8B committed its final answer just 15% to 24% into unconstrained diffusion decoding while half the reasoning region was still masked, producing answer-only outputs on as many as 90% of longer-canvas GSM8K problems. The commitment-order study raised accuracy from 52.8% to 85.2% with frontier-gated decoding while preserving up to fourfold parallelism, and reproduced the mechanism on Dream-7B and MATH-500. Builders attracted to DiffusionGemma’s parallel-token speed should therefore benchmark sampler order alongside throughput: unrestricted parallel commitment can spend the speed advantage by locking an answer before the reasoning exists.