Many diffusion based approaches have been tried for language. I like a diffusion style that goes a bit like this:
* empty string
* Berlin
* Berlin city
* Berlin is a city
* Berlin is a city in Germany
* Berlin is the capital of Germany
* Berlin is the capital of Germany, located in the North East.
* Berlin is the capital. It is located in the North East of Germany.
* Berlin is the populous capital city of Germany. It is in the North East of Germany, by the river Spree.
* Berlin is the capital city of Germany with 3.6 million inhabitants. It is located in the North East of Germany. It's centred on the river Spree.
...
At every step, the diffusion slightly rephrases to add more information content and detail to that of the previous level, by inserting, replacing and deleting tokens.
It is usually more costly to sample from these diffusion style models (which is not the case for diffusion for images, as unlike text, images tend to have a fixed size). But, one might imagine this approach to scale up to generate texts of arbitrary lengths. Just keep on adding more and more detail on your text, and you'll end up with a book or a coherent novel.