Strengths and limitations of diffusion language models
seangoedecke.com
seangoedecke.com
That’s how it does work, but unfortunately denoising the last paragraph requires computing attention scores for every token in that paragraph, which requires checking those tokens against every token in the sequence. So it’s still much less cacheable than the equivalent autoregressive model.
There is quite a bit of evidence diffusion models work better at reasoning because they don't suffer from early token bias.
https://github.com/HKUNLP/diffusion-vs-ar https://arxiv.org/html/2410.14157v3