Yeah for comparison, `tinygrad` takes a little over a second per iteration on my machine.
https://github.com/tinygrad/tinygrad/blob/master/examples/st...
The fastest implementation on my 2060 laptop is AITemplate, being about 2x faster than pure optimized HF diffusers.