This should be much more efficient in theory, right? Why dont we see more leading labs adopt this?
Diffusion text models are cool, but they're functionally much less reliable than autoregressive transformers... and man that's really saying something. Right now most research on them is trying to figure out what complementary systems they need to be reasonably useful.