It’s fast too! I would reckon about 2-3x faster than non-turbo SDXL.
It’s fast too! I would reckon about 2-3x faster than non-turbo SDXL.
Had a test of it and my option is it's an improvement when it comes to following prompts and I do find the images more visually appealing.
Turbo models are decent at low-iteration-decent-results, but not so much at adding fine details to an mostly-done image.
This text was part of the Stability Japan leak (the 20gb VRAM reference was dropped in the release today):
"Stages C and B will be released in two different models. Stage C uses parameters of 1B and 3.6B, and Stage B uses parameters of 700M and 1.5B. However, if you want to minimize your hardware needs, you can also use the 1B parameter version. In Stage B, both give great results, but 1.5 billion is better at reconstructing finer details. Thanks to Stable Cascade's modular approach, the expected amount of VRAM required for inference can be kept at around 20GB, but can be reduced even further by using smaller variations (as mentioned earlier, this (which may reduce the final output quality)."
Maybe running stage C first, unloading it from VRAM, and then do B and A would make it fit in 12 or even 8 GB, but I wonder if the memory transfers would negate any time saving. Might still be worth it if it produces better images though.
Any speed benefits of the 4080 are gonna be worthless the second it has to cycle a model in and out of ram anyway vs the 3090 in image gen.
How is the halo product of a range the "sweet spot"?
I think nVidia are extremely exposed on this front. The RX 7900XTX is also 24GB and under half the price (In UK at least - £800 vs £1,700 for the 4090). It's difficult to get a performance comparison on compute tasks, but I think it's around 70-80% of the 4090 given what I can find. Even a 3090, if you can find one, is £1,500.
The software isn't as stable on AMD hardware, but it does work. I'm running a RX7600 - 8GB myself, and happily doing SDXL. The main problem is that exhausting VRAM causes instability. Exceed it by a lot, and everything is handled fine, but if it's marginal... problems ensue.
The AMD engineers are actively making the experience better, and it may not be long before it's a practical alternative. If/When that happens nVidia will need to slash their prices to sell anything in this sphere, which I can't really see themselves doing.
It's just as likely that AMD will raise prices to compensate.
Granted you could see a supply/demand related increase from retailers if demand spiked, but that's the retailers capitalising.
Because it’s actually a bargain second hand (got another for £650 last week buy it now eBay) and cheap for the benefit it offers for any professional who needs it.
3090 is the iPhone of AI, people should be ecstatic it even exists not complaining about it.
You're aware the 3090 is not the current generation? You can see why I would think you were talking about the 4090?
For example I pulled a (2GB I think, 4 tops) 6870 out of my desktop because it's a beast (in physical size, and power consumption) and I wasn't using it for gaming or anything, figured I'd be fine just with the Intel integrated graphics. But if I wanted to play around with some models locally, it'd be worth putting it back & figuring out how to use it as a secondary card?
Much faster models will come
The iGPU is gfx1036 (RDNA 2).
On my 5900X, so 12 cores, I was able to get SDXL to around 10-15 minutes. I did do a few things to get to that.
1. I used an AMD Zen optimised BLAS library. In particular the AMDBLIS one, although it wasn't that different to the Intel MKL one.
2. I preload the jemalloc library to get better aligned memory allocations.
3. I manually set the number of threads to 12.
This is the start of my ComfyUI CPU invocation script.
export OMP_NUM_THREADS=12
export LD_PRELOAD=/opt/aocl/4.1.0/aocc/lib_LP64/libblis-mt.so:$LD_PRELOAD
export LD_PRELOAD=/usr/lib/libjemalloc.so:$LD_PRELOAD
export MALLOC_CONF="oversize_threshold:1,background_thread:true,metadata_thp:auto,dirty_decay_ms: 60000,muzzy_decay_ms:60000"
Honestly, 12 threads wasn't much better than 8, and more than 12 was detrimental. I was memory bandwidth limited I think, not compute.