Show HN: Realtime LCM tool
twitter.com
twitter.com
AITemplate, stable-fast, or even a torch.compile could get that down to 60ms, I bet. Though I'm not sure the example implementations would work on a non SD architecture.
Hmm, well if you mean torch.compile, y'all should still check out stable-fast, which is claiming ~16ms/iter on a 4090, twice that of torch.compile:
https://github.com/chengzeyi/stable-fast#rtx-4090-512x512-ba...
we still have not been able to replicate those results; also because we want to do it in a distributed way
there's a very cool consistent hashing and probability stuff going on for the GPU routing logic but again, i'll write it after this fire ends similar to a post-mortem.
I'm excited about what H200s or B100s will mean for technology like this one.
there was another video i put on my twitter account but it's already 24 hours old; the UI has changed a bit.
here's the tweet: https://twitter.com/asciidiego/status/1723804366434627611
also, how's infisical doing?