Are y'all running the diffusion in PyTorch eager mode?
AITemplate, stable-fast, or even a torch.compile could get that down to 60ms, I bet. Though I'm not sure the example implementations would work on a non SD architecture.
AITemplate, stable-fast, or even a torch.compile could get that down to 60ms, I bet. Though I'm not sure the example implementations would work on a non SD architecture.
Hmm, well if you mean torch.compile, y'all should still check out stable-fast, which is claiming ~16ms/iter on a 4090, twice that of torch.compile:
https://github.com/chengzeyi/stable-fast#rtx-4090-512x512-ba...
we still have not been able to replicate those results; also because we want to do it in a distributed way