Thank you, I'm not used to reading this kind of research papers but I think I got the gist of it now.
Can this architecture be used to distill models that need fewer timesteps like LCMs or SDXL turbo?
Can this architecture be used to distill models that need fewer timesteps like LCMs or SDXL turbo?