StreamDiffusion: A pipeline-level solution for real-time interactive generation
github.com
github.com
I think that it's possible to get faster than their default timings for a 4090 (I have been able to get 10fps without optimizations with SDXL Turbo and 1 iteration step), but their other improvements like using a Stochastic Similarity Filter to prevent unnecessary generations are good for getting fast results w/out having to pin your GPU at 100% all the time.
I’ve had to git clone everything like a package manager-less serf. Where even is my filesystem?!
On the bright side, one source of "sanity" that I'm finding is to review a collection of daily "hot" publications in AI/ML curated here: https://huggingface.co/papers
Coherent models required dozens of billions of parameters just a year ago and something like Mistral was closer to 70B and reserved for server farms or crazy setups.
Locallama on Reddit has outdated info in from just a few months ago.
Are you an LLM?
> This is how this all will play out in the end right?
Somebody is gonna tell him right? I don't want to be the one to crush such innocence.
EDIT: The `examples\screen` demo almost feels realtime. Says 4 fps on the window but don't what it represents.
EDIT: Denoising in img2img is very low though which means thee returned image is only slightly different from base image.
The slow part for models is loading the model up. But once the model is up, you can send it whatever input that you want.
Parsing and sending the image data just doesn't pass my gut check as to what would be the bottleneck here.
For SD1.5 4090 is able to do ~17it/s without batching and ~90-100it/s with batching.
Although these numbers might be old at this point, I looked at it ~3mo ago.
Their paper is ambiguous unfortunately. Abstract, intro, and conclusion suggests single image by motivating with sequential generation (specifically mentioning metaverse). Experiment section says
> We note that we evaluate the throughput mainly via the average inference time per image through processing 100 images.
That implies batch along with their name Stream Batch...
Looking at the code I'm a bit confused. I'm away from my GPU so can't run. Maybe someone can let me know? This block[0] measures correctly but is using a downloaded image? Then just opens the image in the preprocess? (multi looks identical) This block[1] is using CPU? But running CPU. (there's another like this)
So I'm quite a bit confused tbh.
[0] https://github.com/cumulo-autumn/StreamDiffusion/blob/03e2a7...
[1] https://github.com/cumulo-autumn/StreamDiffusion/blob/03e2a7...
Good job. Give it a try. Look into the server.py of realtime-txt2img to change the model if you want to generate something other than anime. Pointing it to say https://huggingface.co/runwayml/stable-diffusion-v1-5 works fine.
The results are genuinely fast. Not great, but fast. If you change to the SDXL via LCM-LoRA https://huggingface.co/latent-consistency you may get better stuff but that's when it's going to get difficult and you'll start to run into those mysterious crashes I talked about that require, you know, actual work.
my setup: 4090/3990x/CUDA 12.2/debian sid. ymmv.
Stochastic Similarity Filter reduces processing during video input by minimizing conversion operations when there is little change from the previous frame, thereby alleviating GPU processing load, as shown by the red frame in the above GIF
However, a Studio with an M1 Max 64GB is ~13x slower at generative AI with SD1.5 and SDXL than an RTX 4090 24GB at the same cost (~$1,800, refurb) right now.
Does the 4090 have a computer attached to it? It seems like with no computer, the speed would also be 0.
https://support.apple.com/en-us/102363
Five years, and still no solution. And somehow they're spinning memory bandwidth as some sort of prescient act of Apple genius for AI. It's insulting.
Apple made a strange choice with their hardware that effectively pushed our development to Linux and Windows. If Macs didn't make such nice front ends, they almost wouldn't have a place at all.
Mac is not a great place for AI/ML yet. Both the hardware and the software present challenges. It'll take time.
When I was hacking AI stuff on a Macbook, I had a second Framework laptop with EGPU that I SSH'd to.
That said, I think Apple will have some interesting stuff in a year or two (M4 or more likely M5) where they can flex their NPU, Accelerate framework, and unified memory GPU and have it work with more modern requirements.
Time will tell what their software and hardware story is for local inference for generative AI.
Siri (dictation, some assistant stuff, and TTS) runs on device, and I doubt they want to undo that.
I doubt they will do much for training, but maybe a NUMA version of a MacPro with several M4 Ultras will prove me wrong?
Plus two years for software support by the broader ecosystem.
Even Windows, with Cuda + drivers, suffers from less support.
I get a 512x512 5 step image generated in 5 seconds. No refiner, upscaler, or face restoration.
My understanding is that DrawThings hasn’t been optimized for SDXL Turbo and/or pipelined generation yet.
For reference: SDXL Base+Refiner with face restoration at 2k x 2k 50 step image generation takes about 120 seconds.
And this appears to be a local runtime stable diffusion streaming library?
Bruh.
The results aren't even very good. They claim 60x speedup, but compared to what? HuggingFace's Diffusers Autopipeline... a company notorious for buggy code and inefficient pipelines. And that's for naively running the pipeline on every image. Give me a break.
The complexity is in the hardware. Programming has only ever been templating desired machine state. Programmers fell into a religious like state of seeing their more ornate efforts as essential to making a purpose built counting machine count.
The reason I don't have a problem is I see papers as how we researchers communicate with other researchers. But I feel that's not how everyone sees them and there's the aspect that this is how we're judged so incentives get misaligned with actual goal. Idk if the reward hacking is ironic or makes sense because our job is to optimize. But don't let anyone try to convince you that reward (or any cost function) is enough.
ML is crazy right now and people don't see papers as means of researchers communicating to other researchers. You write papers to reviewers. But your reviewers are stochastic so it's hard to write to them because they may or may not be in your niche.
I'll add though that this isn't why journals were created and that CS/ML doesn't typically use journals ({T,J}MLR, PAMI, and a few exist, sure) and instead write to conferences. Fixed dates, zero sum, 1.5 shot setting (1 rebuttal, zero revisions). Journals were created for dissemination of papers, indirectly about communicating to one another, but you know... now we got Arxiv and blogs and websites are sometimes way better just like how papers got better with pictures with computer graphics.