Diffusion Forcing: Next-Token Prediction Meets Full-Sequence Diffusion
boyuan.space
boyuan.space
There are a bunch of neat things you can do with this: in particular, you can firm up parts of the image earlier than others, and thus use it for, say maze solving. They even show it controlling a robot arm moving fruit around, which is pretty wild.
In a way the title undersells the idea - this is a way to do fractional masking, since the masking level is a float - and I think is really a pretty profound and interesting idea.
However, there’s a lot not talked about in this paper; I’d be very curious to see their codebase. How exactly do you set up a maze-following task vs a video extension task? How do you hook up a robot arm to this model, and tell the model what you want done? The architecture itself deserves a significant number of papers / explication.
I assume this is set up so that all tasks are treated as variable horizon, and the current state as a consequence of preceding actions. I agree it would be nice to see the code.
Is the linked codebase enough? I'd be interested to understand what's missing here.
What is the problem you're trying to solve? Are you proposing a new generative model?
> The name "Diffusion Forcing" comes from "teacher forcing" and "diffusion models".