Large Language Diffusion Models
ml-gsai.github.io
ml-gsai.github.io
But at face value, a new architectural approach with the same capacity (8b) trained on a dataset 1/6th the tokens, being competitive with llama3-8b is exciting
"Please recommend me three famous movies"
"The Empire Strikes Back (1980) - Directed by George Lucas"
The paper seems to use EOS padding to create fixed length input/output.
so is there a maximum output length?
In a normal transformer the incremental message is a single token while in LLADA the incremental message would be an additional fixed size block of tokens. If you do inference in the semi-autoregressive strategy outlined in the paper you could start by adding block after block. When a block is fully unmasked it essentially becomes part of the prompt. You would stop the "turn" as soon as there is some sort "end of reply" token decoded in your block. This token could appear anywhere in the block.
Questions: (1) Is there such thing as a "end of reply" token ? (I'm not taking about End of Sentence <EOS> which is obviously different)
(2) What about the junk tokens that appear after this "end of reply" tokens. I could image the unmasking transformer generating some stuff for those tokens.
EOS as end of sentence? IIRC its end of sequence, which may also answer the question IIUC.
So, if we simply wait for <EOS> to be sampled in a block, that could be the place to stop. My answer above should be the same (modulo this correction) -- simply keep obtaining the blocks in a semi-autoregressive fashion.
Does this force the model to encode a high-level answering strategy? (AFAIU, there's no reordering during sampling.) Or does it mean a masking model of a certain size is more prone to making things up that fit the blank space?