Stable Cascade
github.com
github.com
It’s fast too! I would reckon about 2-3x faster than non-turbo SDXL.
For example I pulled a (2GB I think, 4 tops) 6870 out of my desktop because it's a beast (in physical size, and power consumption) and I wasn't using it for gaming or anything, figured I'd be fine just with the Intel integrated graphics. But if I wanted to play around with some models locally, it'd be worth putting it back & figuring out how to use it as a secondary card?
Much faster models will come
The iGPU is gfx1036 (RDNA 2).
On my 5900X, so 12 cores, I was able to get SDXL to around 10-15 minutes. I did do a few things to get to that.
1. I used an AMD Zen optimised BLAS library. In particular the AMDBLIS one, although it wasn't that different to the Intel MKL one.
2. I preload the jemalloc library to get better aligned memory allocations.
3. I manually set the number of threads to 12.
This is the start of my ComfyUI CPU invocation script.
export OMP_NUM_THREADS=12
export LD_PRELOAD=/opt/aocl/4.1.0/aocc/lib_LP64/libblis-mt.so:$LD_PRELOAD
export LD_PRELOAD=/usr/lib/libjemalloc.so:$LD_PRELOAD
export MALLOC_CONF="oversize_threshold:1,background_thread:true,metadata_thp:auto,dirty_decay_ms: 60000,muzzy_decay_ms:60000"
Honestly, 12 threads wasn't much better than 8, and more than 12 was detrimental. I was memory bandwidth limited I think, not compute.This text was part of the Stability Japan leak (the 20gb VRAM reference was dropped in the release today):
"Stages C and B will be released in two different models. Stage C uses parameters of 1B and 3.6B, and Stage B uses parameters of 700M and 1.5B. However, if you want to minimize your hardware needs, you can also use the 1B parameter version. In Stage B, both give great results, but 1.5 billion is better at reconstructing finer details. Thanks to Stable Cascade's modular approach, the expected amount of VRAM required for inference can be kept at around 20GB, but can be reduced even further by using smaller variations (as mentioned earlier, this (which may reduce the final output quality)."
Maybe running stage C first, unloading it from VRAM, and then do B and A would make it fit in 12 or even 8 GB, but I wonder if the memory transfers would negate any time saving. Might still be worth it if it produces better images though.
Any speed benefits of the 4080 are gonna be worthless the second it has to cycle a model in and out of ram anyway vs the 3090 in image gen.
How is the halo product of a range the "sweet spot"?
I think nVidia are extremely exposed on this front. The RX 7900XTX is also 24GB and under half the price (In UK at least - £800 vs £1,700 for the 4090). It's difficult to get a performance comparison on compute tasks, but I think it's around 70-80% of the 4090 given what I can find. Even a 3090, if you can find one, is £1,500.
The software isn't as stable on AMD hardware, but it does work. I'm running a RX7600 - 8GB myself, and happily doing SDXL. The main problem is that exhausting VRAM causes instability. Exceed it by a lot, and everything is handled fine, but if it's marginal... problems ensue.
The AMD engineers are actively making the experience better, and it may not be long before it's a practical alternative. If/When that happens nVidia will need to slash their prices to sell anything in this sphere, which I can't really see themselves doing.
It's just as likely that AMD will raise prices to compensate.
Granted you could see a supply/demand related increase from retailers if demand spiked, but that's the retailers capitalising.
Because it’s actually a bargain second hand (got another for £650 last week buy it now eBay) and cheap for the benefit it offers for any professional who needs it.
3090 is the iPhone of AI, people should be ecstatic it even exists not complaining about it.
You're aware the 3090 is not the current generation? You can see why I would think you were talking about the 4090?
Had a test of it and my option is it's an improvement when it comes to following prompts and I do find the images more visually appealing.
Turbo models are decent at low-iteration-decent-results, but not so much at adding fine details to an mostly-done image.
From what I understand, Stability AI is currently VC funded. It’s bound to burn through tons of money and it’s not clear whether the business model (if any) is sustainable. Perhaps worthy of government funding.
> The company is spending significant amounts of money to grow its business. At the time of its deal with Intel, Stability was spending roughly $8 million a month on bills and payroll and earning a fraction of that in revenue, two of the people familiar with the matter said.
> It made $1.2 million in revenue in August and was on track to make $3 million this month from software and services, according to a post Mostaque wrote on Monday on X, the platform formerly known as Twitter. The post has since been deleted.
https://fortune.com/2023/11/29/stability-ai-sale-intel-ceo-r...
The main reason is probably Mid journey and OpenAi using their tech without any kind of contribution back. AI desperately needs a GPL equivalent…
Why not just the GPL then?
The model weights are shared openly, but the training data used to create these models isn't. This is at least partly because all these models, including OpenAI's, are trained on copyrighted data, so the copyright status of the models themselves is somewhat murky.
In the future we may see models that are 100% trained in the open, but foundational models are currently very expensive to train from scratch. Either prices would need to come down, or enthusiasts will need some way to share radically distributed GPU resources.
There are approaches to get the right type of augmented and generated data to feed these models right, check out our QDAIF paper we worked on for example
Just did some pes2o ablations too which were eh
One thing that’s hard to measure is the knowledge contained in books3. If someone asks about certain books, it won’t be able to give an answer unless the knowledge is there in some form. I’ve often wondered whether scraping the internet is enough rather than training on books directly.
But be careful about relying too much on evals. Ultimately the only benchmark that matters is whether users find the model useful. The clearest test of this would be to train two models side by side, with and without books3, and then ask some people which they prefer.
It’s really tricky to get all of this right. But if there’s more details on the pes2o ablations I’d be curious to see.
For MJ in particular, knowing that they at least used to use Stable Diffusion under the hood, it would not surprise me if the majority of the secret sauce is actually a middle layer that processes the prompt and converts it to one that is better for working with SD. Prompting SD to get output at the MJ quality level takes significantly more tokens, lots of refinement, heavy tweaking of negative prompting, etc. Also a stack of embeddings and LoRAs, though I would place those more in the category of finetuning like you had mentioned.
I've also used it to generate art for deskmats I got printed at https://specterlabs.co/
For commercial stuff I still pay human artists.
I like them quite a bit, and you can get basically any size cut to fit your needs even if they don't directly offer it on the site.
Of course, art is so subjective none of this has any real meaning. MJ routinely blows my mind though and it is very rare something from SD does. The secret MJ sauce is obviously all the human feedback that has gone into the model at this point.
I think AI video will be a different story though. I think that is when comfyUI/SD will destroy MJ because MJ is simply not going to be able to have an economic model with the amount of compute needed.
For fine art MJ destroys everything and it is even close. I say this with using comfyUI/SD all the time.
But you do you and your anime porn.
In fact, anything that involves non-explicitly guided one-shot generation of anything with light/shadow/colors/perspective is entirely out of the question with the current crop, because all models are hallucinating hard and aren't controllable within a single generation. There are attempts at fixing the perspective without explicit guidance, but it's going to be a long way and it's not super relevant to how things are done anyway.
And for fine art, nothing beats a human painter, doing it by throwing prompts at AI mostly misses the point. I'm not even sure what you mean by fine art in this context, actually - surely not generating artsy-looking images from a prompt for fun?
I am not sure if that is still the case.
Given how much VC money is chasing the AI space, this isn't necessarily a bad plan. Give stuff away for free while developing deep expertise, then either figure out something to sell, or pivot to proprietary, or get aquihired by a tech giant.
Instead we've given 10m+ supercomputer hours in grants to all sorts of projects, now we have our grant team in place & there is a huge increase in available funding for folk that can actually build stuff we can tap into.
All the original Stable Diffusion researchers (Robin Rombach, Patrick Esser, Dominik Lorenz, Andreas Blattman) also work for Stability AI.
HN search doesn't seem to agree with me today though and I cannot find the specific comment/s I have in mind, maybe someone else has any luck? This is their user https://news.ycombinator.com/user?id=emadm
We now have top models of every type, sites like www.stableaudio.com, memberships, custom model deals etc so lots of demand
We're the only AI company that can make a model of any type for anyone from scratch & are the most liked / one of the most downloaded on HuggingFace (https://x.com/Jarvis_Data/status/1730394474285572148?s=20, https://x.com/EMostaque/status/1727055672057962634?s=20)
Its going ok, team working hard and shipping good models, the team are accelerating their work on building ComfyUI to bring it all together.
My favourite recent model was CheXagent, I think medical models should be open & will really save lives: https://x.com/Kseniase_/status/1754575702824038717?s=20
Is it legal to use an older snapshot before the license was changed in accordance with the previous MIT license?
Generally courts are more holistic and look at intent, and understand that clerical errors happen. One exception to this is if a business claims it relied on the previous license and invested a bunch of resources as a result.
I believe the timing of commits is pretty important— it would be hard to claim your business made a substantial investment on a pre-announcement repo that was only MIT’ed for a few hours.
Licenses are important. If you are going to expose your code to the world, make sure it has the right license. If you publish your code with the wrong license, you shouldn't be allowed to take it back. Not for an organization of this size that is going to see a new repo cloned thousands of times upon release.
For the same reason you cannot publish a private corporate repo with an MIT license and then have other people claim in “good faith” to be using it.
All they need is to assert that the license was published in error, or that the person publishing it did not have the authority to publish it.
You can’t “magically” make a license stick by putting it in a repo, any more than putting a “name here” sticker on someone’s car and then claiming to own it.
The license file in the repo is simply the notice of the license.
It does not indicate a binding legal agreement.
You of course, can challenge it in court, and ianal, but I assure you, there is president in incorrectly labelled repos removing and changing their licenses.
Show me where they did this. All they have is a commit that says "Update License". An update doesn't imply a correction.
Stability's initial commit had an MIT license, so you can fork that commit and do whatever you want with it. It's MIT licensed.
Now, the tricky part here is that they committed a change to the license that changes it from MIT to proprietary, but they didn't change any code with it. That is definitely invalid, because they cannot license the exact same codebase with two different contradictory licenses. They can only license the changes made to the codebase after the license change. I wouldn't call it "illegal", but it wouldn't stand up in court if they tried to claim that the software is proprietary, because they already distributed it verbatim with an open license.
Why couldn't they? Of course they can. If you are the copyright owner, you can publish/sell your stuff under as many licenses as you like.
That's a marketing opportunity being missed, especially given how crowded the space is now. The HN crowd is more likely to run it themselves when presented with signing up just to test out a single generation.
I guess my zero attempts per minute was too much. You really shouldn't post your product on HN if you aren't prepared for it to work. Reputations are hard to earn, and you're losing people's interest by directing them to a broken product.
[1] https://fal.ai/models/stable-cascade?share=13d35b76-d32f-45c...
Gemini Advanced> Summarize this video: https://www.youtube.com/watch?v=ogJsCPqgFMk
This video is about a new method for training text-to-image diffusion models called Würstchen. The method is significantly more efficient than previous methods, such as Stable Diffusion 1.4, and can achieve similar results with 16 times less training time and compute.The key to Würstchen's efficiency is its use of a two-stage compression process. The first stage uses a VQ-VAE to compress images into a latent space that is 4 times smaller than the latent space used by Stable Diffusion. The second stage uses a diffusion model to further compress the latent space by another factor of 10. This results in a total compression ratio of 40, which is significantly higher than the compression ratio of 8 used by Stable Diffusion.
The compressed latent space allows the text-to-image diffusion model in Würstchen to be much smaller and faster to train than the model in Stable Diffusion. This makes it possible to train Würstchen on a single GPU in just 24,000 GPU hours, while Stable Diffusion 1.4 requires 150,000 GPU hours.
Despite its efficiency, Würstchen is able to generate images that are of comparable quality to those generated by Stable Diffusion. In some cases, Würstchen can even generate images that are of higher quality, such as images with higher resolutions or images that contain more detail.
Overall, Würstchen is a significant advance in the field of text-to-image generation. It makes it possible to train text-to-image models that are more efficient and affordable than ever before. This could lead to a wider range of applications for text-to-image generation, such as creating images for marketing materials, generating illustrations for books, or even creating personalized avatars.
Better coming
ref.: "The model can also understand image embeddings, which makes it possible to generate variations of a given image (left). There was no prompt given here."
This model used to be known as Würstchen v3.
For what it's work, a sort of hybrid CPU GPU approach is giving me workable results with automatic111, about 30-50s to generate a 512x512 image ataround 1-2 it/s.
Not enough for me to want to use local-run txt2img for daily use (why, when Bing/Poe exists and give you high quality image gen for free in seconds), but enough that when I want some granular control or don't want to give my data to a company (i like making AI portraits of myself and friends), I can run locally.
If I could improve performance, I'd love to get into video production
4*16*24*24*4 = 147,456
vs (removing the alpha channel as it's unused here)
3*3*1024*1024 = 9,437,184
Or 1/64 raw size, assuming I haven't fucked up the math/understanding somewhere (very possible at the moment).
I remember setting up the python env for stable diffusion, but then shortly after there were a host of nice GUIs. Are there some popular GUIs that can be used to try out newer models? Similarly, what's the best GUI for some of the older models? Preferably for macos.
Not entirely sure we'll be in the Stable Cascade race quite yet. Since Auto/Comfy aren't really built for businesses, they'll get it incorporated sooner vs later.
Invoke's main focus is building open-source tools for the pros using this for work that are getting disrupted, and non-commercial licenses don't really help the ones that are trying to follow the letter of the license.
Theoretically, since we're just a deployment solution, it might come up with our larger customers who want us to run something they license from Stability, but we've had zero interest on any of the closed-license stuff so far.
ComfyUI is cool but very DIY. You don't get good results unless you wrap your head around all the augmentations and defaults.
No idea if it will support cascade.
There are also a large amount of resources available for it on YouTube, GitHub (https://github.com/comfyanonymous/ComfyUI_examples), reddit (https://old.reddit.com/r/comfyui), CivitAI, Comfy Workflows (https://comfyworkflows.com/), and OpenArt Flow (https://openart.ai/workflows/).
I still use AUTO1111 (https://github.com/AUTOMATIC1111/stable-diffusion-webui) and the recently released and heavily modified fork of AUTO1111 called Forge (https://github.com/lllyasviel/stable-diffusion-webui-forge).
But yeah, the inference speed improvement is mediocre (until I take a look at exactly what computation performed to have more informed opinion on whether it is implementation issue or model issue).
The prompt alignment should be better though. It looks like the model have more parameters to work with text conditioning.
That's an odd comment to place in a thread about an image generation model that is bigger than SDXL. Yes, it works in a smaller latent space, yes its faster in the hardware configuration they've used, but its not smaller.
https://medium.com/@furkangozukara/stable-cascade-prompt-fol...
My Gradio APP even works amazing on 8 GB gpu with CPU offloading
This principle is the idea behind all Stable Diffusion models, this one "just" achieved a much better compression ratio
In which case what are you doing, exactly? Normally you feed it a text prompt instead, which won't compress to the same thing.
There is a model that is trained to compress (very lossy) and decompress the latent, but it's not the main generative model, of course the model doesn't store images in it, you just give the encoder an image and it will encode it and then you can decode it with the decoder and get a very similar image, this encoder and decoder is used during training so that the stage C can work on a compressed latent instead of directly at the pixel level because it's expensive, but the main generative model (stage C) should be able to generate any of the images that were present in the dataset or it fails to do its job. Stages C, B, and A do not store any images.
The B and A stages work like an advanced image decoder, so unless you have something wrong with image decoders in general, I don't see how this could be a problem (a JPEG decoder doesn't store images either, of course).