No Priors
Diffusion will win on inference economics, not on raw intelligence
Stefano Ermon's bet is that autoregressive models have a structural problem nobody can engineer away: generating one token at a time is memory bound and maps badly to GPUs, while diffusion processes many tokens at once and looks at inference like it looks at training. "The bitter lesson is that the more parallel solution is the one that is eventually going to win." Inception's Mercury models now match the speed-optimized tiers from frontier labs on benchmarks while running significantly faster, and voice agent company Open Call moved off Cerebras custom silicon to get comparable speed on ordinary NVIDIA GPUs. He estimates 20 to 30 percent of workloads today are latency bound, and flags a second advantage: because diffusion refines coarse to fine, you can score and steer a partial output instead of waiting for the finished object.
- #hardware
- #open-source
- #products
