> For distilled StableDiffusion 2 which requires 1 to 4 iterations instead of 50, the same M2 device should generate an image in <<1 second
> For distilled StableDiffusion 2 which requires 1 to 4 iterations instead of 50, the same M2 device should generate an image in <<1 second
They have some benchmarks on the github repo: https://github.com/apple/ml-stable-diffusion
For reference, previously I was getting about <3 minutes for 50 iterations on my Macbook Air M1. I haven't yet tried Apple's implementation but it looks like a huge improvement. It might take it from "possible" to "usable".
So sounds plausible that the m1 can reach the same level in some use cases with the right optimizations.
Mac Studio with M1 Ultra gets 3.3 iters/sec for me.
MacBook Pro M1 Max gets 2.8 iters/sec for me.
But also yes, it's gotta be expensive to host these models and I'm not sure where all these subsidies are coming from. I expect that we'll eventually see these things transition to more paid services.
https://explosion.ai/blog/metal-performance-shaders
However the SoC only uses 31W when posting that performance.
And the posted benchmarks for the M2 Macbook Air make me consider 'upgrading' to an Air.
Macbook Airs (way back when) felt sluggish. The MBA M1 changed that, it was "fine". These M2s are unexpectedly responsive on an ongoing basis.
The MacBook Pro M1 Max is great (would be fantastic except they lost a Thunderbolt port in favor of legacy HDMI and memory card jacks), but you expect that machine to be responsive, so it's less surprising.
The Studio Ultra, though, never slows down for anything.
Still, if the Air could drive two external screens instead of one, I'd "downgrade" from the Max.
I've since moved to the M2 air, and it is noticeably faster than M1, but it isn't the huge leap from last gen intel that the M1 was. But the hardware itself feels way better.
* Air built-in display
* 2K display connected via USB-C -> DisplayPort adapter
* Two more 2K displays of same model via DisplayLink connected via USB hub
For all practical means it's almost impossible to see any DisplayLink compression artifacts even in most of games.
PS: Each adapter cost me $40:
Been nervous to dip into it, given the architecture change and last year's challenges with display link docks.
// UPDATE: Oops, looking at the product, I see I should have specified: 4K screens or higher. About half our desks are 2 x 4K, about half 2 x 5K, except the Air M1 folks who are 1 x 5K.
For higher resolution some other solution is required.
Dall-e et. al will still be able to bandwagon off of all the free ecosystem being built around the $10M SD1.4 model that is showing what is possible.
E.g. Dall-e could go straight to Hollywood if their model training works better than SD’s. The toolsets will work
Maybe a dumb question but can the old model still be run?
SD2 wasn’t “neutered”, the piece of it from OpenAI that knew a lot of artist names but wasn’t reproduceable was replaced with a new one from Stability that doesn’t. You can fine-tune anything you want back in.
The "in the style of Greg Rutkowski" prompts from SD1 though, IIRC, were thought to be proof it was reproducing the training set. But it actually only saw ~27 images of his, and the rest was residual biases from CLIP.
https://mezha.media/en/2022/10/06/google-is-working-on-image...
Give it some time and SD will be able to do the same.
What's cool about the era in which we live is if you look at high-performance graphics for games or simulations, for instance, it may in fact be faster to a the model to "enhance" a low-resolution frame rather than trying to render it fully on the machine.
ex. AMD's FSR vs NVIDIA DLSS
- AMD FSR (Fidelity FX Super Resolution): https://www.amd.com/en/technologies/fidelityfx-super-resolut...
- NVIDIA DLSS (Deep Learning Super Sampling): jhttps://www.nvidia.com/en-us/geforce/technologies/dlss/
AMD's approach renders the game at a crummy, low-detail resolution then each frame uses "upscales"
Both FSR and DLSS aim to improve frames-per-second in games by rendering them below your monitor’s native resolution, then upscaling them to make up the difference in sharpness. Currently, FSR uses spatial upscaling, meaning it only applies its upscaling algorithm to one frame at a time. Temporal upscalers, like DLSS, can compare multiple frames at once, to reconstruct a more finely-detailed image that both more closely resembles native res and can better handle motion. DLSS specifically uses the machine learning capabilities of GeForce RTX graphics cards to process all that data in (more or less) real time.
Video is really a series of frames, the framerate for film/human could get away with 24 frames/second-- ~40ms/image for real-time.
What's cool about the era in which we live is if you look at high-performance graphics for games or simulations, it may in fact be faster to run the model on each frame to "enhance" a low-resolution frame rather than trying to render it fully on the machine.
ex. AMD's FSR vs NVIDIA DLSS
- AMD FSR (Fidelity FX Super Resolution): https://www.amd.com/en/technologies/fidelityfx-super-resolut...
- NVIDIA DLSS (Deep Learning Super Sampling): https://www.nvidia.com/en-us/geforce/technologies/dlss/
AMD's approach renders the game at a crummy, low-detail resolution then use "spatial upscaling" to enhance the images one frame at a time.
NVIDIA DLSS uses "temporal upscaling" to pass over multiple frames and uses other capabilities exclusive to Nvidia's cards to stitch together the frames.
This is a different challenge than generating the content from scratch
I don't think this is possible in real-time yet, but someone put a filter trained on the German country side to produce photorealistic Grand Theft Auto driving gameplay:
https://www.youtube.com/watch?v=P1IcaBn3ej0
Notice the mountains in the background go from Southern California brown to lush green
https://www.rockpapershotgun.com/amd-fsr-20-is-a-more-demand....
See deforum[1] and andreasjansson‘s stable-diffusion-animation[2]
[1]: https://deforum.github.io/
[2]: https://replicate.com/andreasjansson/stable-diffusion-animat...