Running Stable Diffusion XL 1.0 in 298MB of RAM
github.com
github.com
"OnnxStream can consume even 55x less memory than OnnxRuntime while being only 0.5-2x slower"
The trade-off between (V)RAM use and inference time sounds like it could be advantageous in some scenarios, and not just when RAM is constrained like in the RPi case.
I actually wonder if this weight unloading approach can be used to handle larger batch sizes in the same amount of RAM, in effect increasing throughput massively at the cost of latency.
I assume they meant to say "1.5-2x slower".
And explain to me why it isn't you?
38.157% of informally provided statistics are made up on the spot under the assumption nobody will actually check.
I’ve been doing capacity planning lately and the whole deal with how we are bad at fractions came up again. If you have to spool up 10% more servers and then cut costs by 10%, you’re still slightly ahead. If you cut 10% and then later another 10% you have cut 19% of the original, not 20%.
When we say "faster" or "slower" what we usually mean is that we add/remove the percentage to the original amount, which is often cause of misunderstanding.
"Y is 10% faster than X" means that Y goes at 110% the speed of X
"Y is 10% slower than X" means that Y goes at 90% the speed of X
In particular, "Y is N% slower than X" doesn't mean that "X is %N faster than Y" ! (110% of 90% is not 100%)
For example "Y is 100% slower than X" doesn't mean that "X is double as fast as Y", but that Y is not moving at all.
and "Y is 200% slower than X" means... that Y goes in the other direction ? (Maybe back in time, in this case ?)
I don't like how it was worded by the author. But all you've done is essentially invert the wording while making the math MORE difficult in the process.
50% to 200% is 0.5x slower to 2x slower.
People seem to be confusing “% slower/more time” vs “% of current time”
I think something like “0.5x slower” is just conventionally not used, even if it’s understandable. It’s a difference in common usage between percentages and x-factors. One reason might be that x-factor implies multiplication, not addition; that’s what the “x” stands for. X-factors are typically used to say something like the ‘the new run time was 1.7x the old one’. Whether it was slower or faster is implied by whether the x-factor is below or above 1.0. Because x-factors are commonly used as multipliers, and not commonly used to say “1.5x more than” (which actually means 2.5x), it’s pretty easy for people to misunderstand when someone says “0.5x slower” because it looks like an x-factor.
Now, the same argument could apply to percentages. A percentage is also a factor. But in actual usage, “50% more” is common and “0.5x more” is not; and “100x faster” is common (usually to mean 100x not 101x) while “10000% faster” is not common at all. So language is inconsistent. ;)
All that said, using an x-factor as a pure factor, and not a multiply-add, is less confusing and more clear. Saying “The new runtime is 1.5 times the old one” leaves no room for error, where “The new runtime is 90% slower than the old one” is actually pretty easy to miscalculate, easy to mistake, and easy to misinterpret. The percentage-add is also asymmetric: 90% slower means 0.1x, while 90% faster means 1.9x. Stating a metric as percentage-add makes sense for small percentages, and makes more sense for add than subtract once the numbers are double-digit and larger.
I think they are correct but it is easy enough to misinterpret that it is not a good way to phrase things.
Is it as fast as the original or does it take twice as long?
People should just use duration instead of speed as you did at the end: "takes twice as long", "takes 1/3 of the time..."
I have played a fair number of incremental games and this quickly became a pet peeve of mine. So many will say things like "2x more" and it will actually be "2x as much". Fortunately, I don't recall any which actually switch between the meanings but it's so commonly a guessing game until I figure it out.
I'm the author.
I have never questioned the clarity of that sentence, at least until today :-)
By "0.5x" I mean "0.5 times or 0.5 multiplied by the reference time" where "reference time" is the inference time of OnnxRuntime. So I'm actually meaning "50%".
I think the expression "0.5x slower", taken by itself, could be misleading, but in the context of the original sentence it becomes clearer ("while being only 0.5-2x slower", the "only" is important here!!!).
But I think the general context of that sentence defines its meaning. I am referring to the fact that under no circumstances could my project, with its premises, be in any scenario even a single millisecond faster than OnnxRuntime. Then in the second paragraph of the README the goal of the project is stated, which is precisely to trade off inference time for RAM usage! Obviously combined with the fact that the performance data is clearly reported and that I repeat several times that the generation of a single image takes hours or even dozens of hours.
However, given the possible misunderstanding, I will correct the sentence in the next few days.
Do you mean it increases the runtime by 50%-200%?
In terms of throughput it is a decrease by 0.33-0.67x to 0.67-0.33x. That this are the same numbers in reverse order is of course just a coincidence, would the runtime have increased by only 0.2-0.5x to 1.2-1.5x, then the throughput would have decreased by 0.17-0.33x to 0.83-0.67x.
Since inference is generally memory bandwidth bound once you reach the level of 'does this model even fit in the given system', I'd imagine that this technique wouldn't help much for greater throughpit via larger batch sizes. Just one instance is probably already saturating the memory controller.
Maybe it'd help on the training side though?
That was about when I decided I needed other hobbies. Right before that happened some brilliant soul put out a tool that would render your scene in OpenGL so you could look at it first. I don't think that would run on your Amiga but it (barely) ran on my machine.
And the end result, a single image of a mediocre scene by a poor artist with amazing caustics. I won't be quitting my day job.
https://github.com/MoonRide303/Fooocus-MRE
For base SD 1.5, I use Volta, because its fast: https://github.com/VoltaML/voltaML-fast-stable-diffusion/com...
Really good SD 1.5 image quality comes from gratuitous use of finetunes, LORAs, controlnet and other augmentations. So you can, say, trace a base image for structure, specify prompting in certain areas of the image and so on. InvokeAI is actually quite feature packed, and has lots of these augmentations hidden in the nodes UI, but Volta and other UIs also expose them more directly.
Still, even with it turned off, the quality is quite remarkable.
Ip adapter uses an image to guide denoising.
Fooocus and MJ take a prompt and expand it in a variety of ways (eg a language model or more simplistic text manipulation). The actual prompt that creates the conditioning is not what you typed in. That’s what I mean by prompt massaging
You’ll need more time and memory compared to Invoke or an Nvidia graphics card, but it’s not that bad: 1-2 s/it for an image in standard 512x768px quality, 14-20 s/it for an image in high 1024x1536px quality (Hires Fix).
Generally the trade off is that any of the impressive finetuned models are far less generalizable then the default weights, but in practice this is not a big deal and the results can be a substantial improvement.
That is big though
And now we're going to put a screen on the wall that we don't even look at 99% of the time?
I think it might be friendlier in some aspects than fetching an image from a server running the big the models.
And you don't have to worry about service disruptions or api keys
An e-ink display doing it should only use energy when refreshes. And you could minimize refreshes to once a day, week, etc
Less friendly than a photo of course.
In all seriousness I can give a brief overview:
- I'll probably offload the image generation to the 5 year old intel nuc I already have as a home automation server, comfyUI in CPU mode takes 20-30 mins for a generation. Ideally it's all self contained on the Pi but that might be beyond me, skill wise.
- prompts are composed by taking time of day, season, special occasions (birthdays, xmas etc); adding random subjects from a long manually curated list; then asking gpt4 to creatively remix the prompt for variety
- i have an inky impression 7.3 inch 7 color eink display and a raspberry zero stuck onto it. Right now it'll simply download new images from the NUC every once in a while
- i like wood and i dislike the jagged 3d printer aesthetic so I'll create a frame from laser cut plywood by designing some stackable svg shapes in inkscape and sending those to a laser cutter
It works right now, functionally.
Considering that I'm painstakingly writing this on a phone with a sleeping 3 week old baby on my chest it'll be while before i have the energy to make it look like something you'd hang on your wall
https://hackaday.com/2023/09/19/e-paper-news-feed-illustrate...
I'd assume the pc would still likely win.
Something like a Pi 4 or 5 may be a better benchmark than the Zero 2 as I get the impression its been used more for the challenge than practicality.
Verily, the era is nigh wherein even lamps and toasters shall brim with surpassing sagacity.
After exposure to this field for many years, the last decade was stunning.
I say “was”, because the speedup in the last 6-18 months has been another thing altogether.
I am not concerned with what we will be able to two years hence, but with how much faster progress will be. And then again, and again.
Let's make a startup!
Or same can be said about media "piracy". Or ransomwares.
States have forever regulated things that are not possible to enforce purely technically.
Piracy most often isn't treated as a criminal matter, but a civil one - few countries punish piracy severely, but companies are allowed to sue the pirate.
I agree with OP in principle - regulating generative AI use would be way harder than piracy or whatever, especially since all of it can be done purely locally and millions of people already have the software downloaded. And that's not getting into the reasoning behind a ban - piracy and similar "digital crimes" are banned because they directly harm someone, while someone launching Stable Diffusion on their PC doesn't do much of anything.
UNCLOS, Part VII, Section 1, Article 100 https://www.un.org/depts/los/convention_agreements/texts/unc...
>> Duty to cooperate in the repression of piracy
>> All States shall cooperate to the fullest possible extent in the repression of piracy on the high seas or in any other place outside the jurisdiction of any State.
We could have just added "private computer" to the definition of piracy, and it largely would have applied.
>> Definition of piracy
>> Piracy consists of any of the following acts:
>> (a) any illegal acts of violence or detention, or any act of depredation, committed for private ends by the crew or the passengers of a private ship or a private aircraft, and directed [...] on the high seas, against another ship or aircraft, or against persons or property on board such ship or aircraft;
No sane person could ever implement anything like this. This is like saying that we could "just" add the word "digital" to the laws prohibiting murder to make playing GTA illegal.
Mostly committed by private citizens in pursuit of profit
That all nations of the world have an interest in suppressing to encourage free trade that economically benefits them
But which some countries at various times have a geopolitical interest in supporting
... you're right, they have no logical or legal connections at all.
The point stands - no jurisdiction that I know of treats digital piracy similarly to naval piracy, and there is no strong argument in favor of doing so.
The canonical eBay/PayPal fraud from eastern Europe example?
> most individuals pirate for personal needs, not profit.
But most piracy is done by individuals in pursuit of profit, not for personal need.
your piracy example is better. consider that it's the rise of more convenient options (netflix and spotify) not some effective policy that curtailed the prevalence of piracy.
The turning point was earlier than Netflix or Spotify – it was the iTunes Store. It was such a dramatic shift, people labelled Steve Jobs as “the man who persuaded the world to pay for content”.
https://www.theguardian.com/media/organgrinder/2011/aug/28/s...
In fact, a low clearance rate can be evidence of trying to regulate far beyond one's capacity to consistently enforce; if you weren't trying to regulate very hard, it would be much easier to have a high clearance rate for violations of what regulations you do have.
It’s too generic. There are too many of them.
How many regular people would risk owning turning-complete devices that can run unauthorized software if it would net you jail time if caught? Lots of countries are already itching towards banning VPN, corpo needs be damned.
Especially now that the iPhone has shown having a device that can only run approved legal software covers a lot of people's everyday needs.
Theoretically, you can have a totally owned device managed by Big Brother, yet generate AI smut with a general purpose CPU built in PowerPoint.
How do you possibly regulate that?
The government could send an order to the software developer to patch out that turning completeness, and ban the software if it's not complied.
I get what you mean, it's never possible to 100% limit things. But if you limit things 98% so that the general public does not have access that's more than enough for authoritarian purposes.
The question in my head is whether the failures in their approaches are due to a flaw in the implementation (in which case it's practically possible to do what they're trying to do although they haven't figured out a way to do it), or whether it's fundamentally impossible. With DRM and content, there's always the analog hole, and if you have physical control over the device, there's always a way to crack the software and the hardware if need be. My questions are whether:
a) this is a workable analogy (I think it's imperfect because Gen AI and DRM are kinda different beasts)
b) even if it was, is there real way to limit Gen AI at a hardware level (I think that's also hard because as long as you can do hardware accelerated matmul it's basically opening up the equivalent of the analog hole towards semi-turing completeness which is also hardware accelerated)
I imagine someone has thought through this more deeply than me and would be curious what they think.
[1] https://en.wikipedia.org/wiki/High-bandwidth_Digital_Content...
[2] https://techaeris.com/2018/02/18/microsoft-uwp-protection-cr...
Netflix for example can implement any DRM tech they want -- ultimately they're putting a picture on my screen, and it's impossible to stop me from extracting it.
https://frame.work/ and the https://mntre.com/ MNT Reform: Exist
Furthermore, it also would mean that I would not be able to bring any personal computers with me when I travel to other countries. I like to travel, and I like to bring my computers when I do.
Next, it would also be dangerous to try to buy computers locally within the borders of the country. The seller might be an informant of the police, or even a LEO doing a sting operation.
And then next you have to worry about the computers you already have. If you decide to keep the computers that you had since before, after it is made illegal to own them, you will have problems even if you keep them hidden and only use them at home. Other people know about your computers. Some of those people will definitely tip off the authorities about the fact that you are known to have computers.
Let’s hope it never goes as far like this :(
What country outside of North Korea has banned the ownership of general purpose computers, or even considered/tried to?
Especially because these tools are so popular outside of the developer community, I think it's worth really beating into peoples minds that without open source AI would be in a much worse place overall.
The original requirement for these is 16GB of RAM, which can be had for less than $20. They run much faster on a GPU, which can be had for less than $200. Millions of ordinary people already have both of these things.
That said, I don't think blanket regulation is all that likely anyhow.
I know it was a bit of a funny hyperbolic example, but you'd need to shrink this down way further to run it on a PS2.
I do not mind waiting for my images, but i still have more to offer than a raspi
This would deter real-time or near real-time applications where latency is a critical factor.
Also, the confusion over the phrase "0.5-2x slower" highlights a possible lack of clarity in communication within the community, which would hinder the accurate assessment and adoption of such optimizations in practice.
For example:
> Also, the confusion over the phrase "0.5-2x slower" highlights a possible lack of clarity in communication within the community, which would hinder the accurate assessment and adoption of such optimizations in practice.
Maybe instead:
> The phrase "0.5-2x slower" is confusing. You might get more adoption if the language was more clear.
(ducks)
So pointless. I love it