No Priors
Diffusion beats autoregression at inference because GPUs hate sequential work
Autoregressive generation is fundamentally memory bound: you cannot produce the tenth token before the ninth, so you spend most of your time shuttling weights around and doing very little arithmetic. Diffusion models process many tokens at once, which makes their inference workload look like their training workload, the thing GPUs are actually good at. Inception's Mercury models match Haiku, Flash, and mini/nano class models on benchmarks while running significantly faster, and one voice customer dropped Cerebras custom silicon for them because software parallelism got the same speed on ordinary NVIDIA GPUs. "The bitter lesson is that the more parallel solution is the one that is eventually going to win."
- #research
- #hardware
- #products
