TREAD: Token Routing for Efficient Architecture-Agnostic Diffusion Training
arxiv.org
arxiv.org
That process means they may require a hundred or more training iterations on a single image. I haven't digested the paper, but it sounds like they are proposing something conceptually similar to skip layers (but significantly more involved).
If so, what are the DiT specific changes that needed to be made?