Generate images in one second on your Mac using a latent consistency model
replicate.com
replicate.com
Edit: it seems the "per second" requires the `--continuous` flag to bypass the initial startup time. With that, I'm now seeing the ~1 second per image time (if initial startup time is ignored).
It redownloads 4 gigs worth of stuff every execution. Can't you have the script save, and check if its there, then download it or am I doing something wrong?
With 5 iterations the quality is...not good. It looks just like Stable Diffusion with low iteration count. Maybe there is some magic that kicks in if you have a more powerful Mac?
I know what I'll be doing this weekend... generating artwork for my 9 yo kid's video game in Game Maker Studio!
Does anyone know any quick hacks to the python code to sequentially prompt the user for input without purging the model from memory?
A few minutes? I have to download at least 5GiB of data to get this running.
https://github.com/replicate/latent-consistency-model/commit...
Follow the instructions. Before actually running the command to generate an image.
Open up main.py Change line 17 to model.to(torch_device="cpu", torch_dtype=torch.float32).to('cpu:0')
Basically change the backend from mps to cpu
* after `pip install -r requirements.txt` do `pip3 install torch torchvision torchaudio xformers --index-url https://download.pytorch.org/whl/cu121`
* on line 17 of main.py change torch.float32 to torch.float16 and change mps:0 to cuda:0
* add a new line after 17 `model.enable_xformers_memory_efficient_attention()`
The xFormers stuff is optional, but it should make it a bit faster. For me this got it generating images in less than second [00:00<00:00, 9.43it/s] and used 4.6GB of VRAM.
But you might be able to generate at 15fps and interpolate between them or something.
Generation takes 20-40 seconds, when using "--continuous" it takes 20-40 seconds once and then keeps generating every 3-5 seconds.
I'd like do to some comparison testing. The model in the post is fast but results are hit or miss for quality.
Its not fast, but its SOTA local quality as far as I know, and I've tried many UIs and augmentations.
Also, maybe it will run better if you grab Pytorch 2.1 or nightly.
This on an M2 Max 32gb
That model doesn't work well at 1024x1024 anyway without some augmentations. You want this instead: https://huggingface.co/segmind/SSD-1B
https://github.com/simple10/ai-image-generator/blob/main/exa...
But thats the point where regular diffusion (with the UniPC scheduler and FreeU) overtakes this in terms of quality.
edit: I tried it out by copying this pipeline file locally and then disabling the safety checker. https://raw.githubusercontent.com/huggingface/diffusers/main...
On my M1 macbook, did a test of 10 images, including the one-off loading time. With checker: 10.51s, without safety checker: 9.48s. So not that big of a hit.
Doing a search for "nsfw" in all subdirectories seems to turn up all the files you need to edit.
On my Desktop PC with a 4090 in I was getting speeds of 0.2 to 0.3 seconds for reasonably acceptable quality settings so I would expect 0.5s or so on the laptop.
What Apple are ahead on is doing this on a fanless laptop that doesn't hit internal temperatures of triple digits.
You also forgot the bit where Apple are ahead of doing it on a laptop that can achieve it without needing to be tethered to a power socket to achieve the performance.
Kind of sad that a huge anti-competitive, trillion dollar company is the one offering it. Especially given their stances around user freedom.
I'd much rather innovation be distributed. The goal posts should be moved to a point everyone is pushing towards the next thing. Having Apple be the only game in town is unhealthy.
I think you could pull this off on a Asus G14 in an ultra power saver mode, with the fans off or running inaudibly. The cooling is so beefy they will actually work fanless if you throttle everything down and mostly keep the GPU asleep.
The M chips could certainly sustain image generation better without a fan.