I'm struggling to understand where the gains are coming from. What is the intuition for why DiT training was so inefficient?
That process means they may require a hundred or more training iterations on a single image. I haven't digested the paper, but it sounds like they are proposing something conceptually similar to skip layers (but significantly more involved).