Stable Diffusion WebGPU demo
islamov.ai
islamov.ai
Even more impressively, they followed up with support for several Large Language Models: https://webllm.mlc.ai/
I've since moved away from ONNX and to a more GGML style.
Being bound to ONNX means moving at a slower velocity - the field moves so fast that you need complete control.
Interesting what browsers have become. The web ate the operating systems.
It'll be the same size download and use about the same RAM if you download it and run it directly without using a browser.
However, the web also has a terrible bloat/legacy issue it refuses to deal with. So sooner or later, a new minimal platform will grow from within it, the way the web started in the 90s as the then humble browser. And the web will be replaced.
First, it's explicitly designed to be easy to port existing software to. Like C libraries, say. Sounds good right? Well, it's not designed as a platform that arose from the needs of the users, like the web did. But the need of developers porting software, who previously compiled C libraries to JS gibberish and had to deal with garbage collection bottlenecks etc. This seems fairly narrow. WASM has almost no contact surface with the rest of the browser, aside from raw compute. It can't access DOM, or GPU or anything (last I checked).
Second, for reasons of safety, they eliminated jumps from the code. Instead it's structured, like a high-level language. It has an explicit call stack and everything. Which is great, except this is a complete mismatch for all new generation languages like Go, Rust etc. which focus heavily on concurrency, coroutines, generators, async functions etc. Translating code like that to WASM requires workarounds with significant performance penalty.
For this one, as long as my browser supports WebGPU (which will be widely supported soon) and I have the system resources, it will run. Barely any technical knowledge needed, doesn't matter what OS or brand of GPU I have. Isn't that really cool? It reduces both technical and knowledge based barriers to entry. Why do people criticize this so strongly?
In a similar vein, you can find plenty of comments on HN faulting the massive proliferation of smartphones among the general population throughout 2010s for "ruining" web, software ecosystems, application paradigms, etc. There are plenty of things one could potentially criticize smartphones for, and some of that criticism indeed has merit. But this specific point about "ruining" things feels almost like a different version of the same argument above - niche things becoming widely adopted by the masses and "ruining" their "cool kids club."
Another similar example from an entirely unrelated domain - comic books and their explosion in popularity after Marvel movies repeatedly killing it in the box office. I don't even like Marvel movies, barely watched any of them, but the elitism around hating things becoming more popular is just silly.
The objection (more surprise than objection) is that web browsers are supposed to be sandboxed environments. They are not supposed to be able to do things that negatively impact system performance. It is surprising you can do things involving multi-gb of ram in a web browser. It has nothing to do with what you are using that ram for or if its cool or not.
I dont think anybody has an objection to making it easier to run stable diffusion and i think the only way you could come to that conclusion is intentionally misunterpreting people's comments.
I agree with the sandboxing model, but it is orthogonal to WebGPU and impacting system performance. Sandboxing is about making the environment hermetic (for security purposes and such), not about full hardware bandwidth isolation.
First, there is no way for web browsers to have no system performance impact whatsoever. Browsers already use hardware acceleration (which you can disable, thus alleviating your WebGPU concerns as well), your RAM, and your CPU.
Second, afaik WebGPU has limits on how much of your GPU resources it is allowed to use (for the exact purpose of limiting system performance impact).
Java's real success was ( and still is ) on the server - powering a whole generation of internet applications, and creating a cross vendor ecosystem that stopped MS leveraging it's client dominance to take over the server space as well.
I don't believe Unix/Linux would have survived the Windows server onslaught without Java on the backend and the web on the front.
But it was never Java's "original premise", which is what the comment you are replying to was about. According to their (very heavy-handed) marketing at the time, Java was supposed to be for native desktop applications and for "applets". But yeah, in the many years it took for those promises to truly become hollow, Java carved out a surprisingly robust niche for itself on the enterprise server.
Also, I am skeptical of this last sentence of yours. The thing that resisted the Windows server onslaught, broadly, was the wide range of free-as-in-speech-and-as-in-beer backend technologies, like Perl, PHP, Python, Postgres, and some other things that start with "P", as well as, yeah, Java. Java played a role, but it was just one of many.
On the desktop we had Swing which was OKish to build GUI apps (albeit still underneath Borland) and then totally lost sight of desktop with JavaFX that was created without hearing the community and then abandoned, also refusing to improve Swing. Quite a pity.
This is technically not true, as far as I know. That whole idea of Java ME, different "profiles", all that stuff happens in roughly 1998, which is definitely not "from the beginning". Though, looking it up now, apparently the Java Card stuff gets started a little earlier than that (which I didn't know/notice at the time, probably because it apparently wasn't initially a Sun initiative and so I'm guessing Sun's self-promotion didn't mention it in the really early days).
But depending on what your point is, maybe my first paragraph is merely a technical quibble, not a substantive disagreement. Maybe your point is that Java's success has been, in part, due to its ubiquity in small-but-not-tiny devices like "feature phones". Fair enough, I guess, and if that's your point then it doesn't really matter if it was truly "from the beginning", or just "one of the earliest pivots" (which I think is more accurate).
Myself, my point is that DrScientist's reply to quickthrower2 is, as a reply, just straight-up wrong wrong wrong. Java's original premise was twofold: web applets, and desktop apps that didn't need to be maximum performance (note that Swing was also not the original Java GUI toolkit, I've forgotten the name of the thing that preceded it, Swing was certainly much better). Building servers was NOT part of Java's original premise. And quickthrower2 is right: the web ate that original premise. Java had to pivot to live, and did.
I'm getting too pedantic here, but the historical revisionism is winding me up.
This was common knowledge on that decade.
From memory don't recall Java being focused on server-side much later until the 2000s with Tomcat and JBoss making a lot of stride, can't say I was fan of either. Maybe that is the time when your person first saw Java trying to compete for whatever space was left of web to take. I'm failing to have the impression AWT was ever relevant, that's why it wasn't even mentioned as everyone seemed to be using only Swing except for some god-awful projects in the gov domain.
For embedded developers (phones, smartcards, electronic devices, ...) it was well-established since the early days because IMHO was _easy_ to use/deploy/maintain by comparison to other options. Even looking at the options available today, it is still on the top albeit C++ making quite a fantastic comeback with Arduino albeit continuing to be a pain in the rear to debug.
https://en.wikipedia.org/wiki/Oak_(programming_language)
No revisionism - you aren't just looking back far enough into history.
The original Java GUI toolkit was called AWT.
See https://wiki.c2.com/?TheStoryOfAwt for some interesting history - as you can see from the story - the web was a pivot, not the original intention.
Exactly.
The biggest deployment of Java was/is on Java smart cards - billions of devices every year.
Or java was too heavy for the computers at the time to get people to use "applets" for everyday things (i.e. go to a new website a do a thing on it)
Flash et al also failed to catch for long.
The web browser's success might have something to do with neverending feature creep as opposed to "this can do everything but as such it's broken and vulnerable".
Memory leaks abounded in particular.
Life cycle management was difficult as well.
Note the interface had to be implemented in each and every browser separately - compounding the problem - for applets to be viable it had to work on all the major browsers well.
Not blaming the people who worked on it - I suspect the origin design was put together in a rush, and the work under-resourced, and it required coordination of multiple parties.
Because if it's the web, Google sees it. And if everything is the web, then Google sees everything.
(I remember several awesome hobby OS projects ported KHTML to get a really good browser back in those days. It was a really solid and portable codebase and much tidier than Firefox.)
Google sees everything that is public and everything that uses their ad network, including data from apps that don't use the web at all.
Can't remember the last time I used Windows for anything more than launching Chrome and Steam...
I guess you would be better off with SteamOS then.
If not for DirectX and Windows-only games, I'd totally ditch it. Maybe when Proton gets there.
"Hey this game wants to use your motion controls and USB gamepad." Okay sure.
Which is exactly what resources are for, when eating is performed correctly :)
Though, they might eat the desktop environment or UIs (in general).
Darn, guess I'll have to wait for stuff to land in Firefox.
How are others getting it working?!
From someone else's comment, this one works fine: https://websd.mlc.ai/#text-to-image-generation-demo
If you can't run it, here's how the output looks with default settings https://i.imgur.com/WCQc8hO.png
> You need latest Chrome with "Experimental WebAssembly" and "Experimental WebAssembly JavaScript Promise Integration (JSPI)" flags enabled!
Now I'm wondering whether the top message goes away once the flags are enabled?
Yes. I thought it won't be good if it would download 3.5gb once you open the page.
>Now I'm wondering whether the top message goes away once the flags are enabled?
No, I haven't added any checks for that (and I'm not sure how the first one can be properly checked), so it's just an info bar. Which is, eventually, misleading.
https://github.com/WebAssembly/js-promise-integration/blob/m...
Implementing a new feature in WebAssembly is a bit more complex due to its execution model and security constraints. I expect it's also just the case that a lot of these new WASM features are very complex - promise integration is super nontrivial to get right, so are WebAssembly GC and SIMD.
Anything beyond those use cases it is really meh, specially given how clunky compiling and debugin WASM code tends to be.
Then we have all those startups trying to reivent bytecode executable formats in the server, as if it wasn't something that has been done every couple of years since late 1950's.
Right but it doesn't right now? Like you can't just write arbitrary code as you would with a Java plugin, or a PNaCL C++ plugin. Wasm is extremely difficult to use for those use cases.
> Then we have all those startups trying to reivent bytecode executable formats in the server, as if it wasn't something that has been done every couple of years since late 1950's.
Yes, because people really want this and the solutions have all been fraught with security issues historically.
"Everything Old is New Again: Binary Security of WebAssembly"
https://www.usenix.org/conference/usenixsecurity20/presentat...
"Swivel: Hardening WebAssembly against Spectre"
https://www.usenix.org/conference/usenixsecurity21/presentat...
Notably, the first paper is about exploitation of webassembly processes. That's valuable but the flaws of previous systems wasn't that the programs in those systems were exploitable but that the virtual machines were. Some of this was due to the fact that the underlying virtual machines, like the JVM, were de-facto unconstrained and the web use case attempted to add constraitns on after the fact; obviously webassembly has been designed differently.
I hope wasm sees more mitigations, but I also expect that wasm is going to be a target primarily for memory safe languages where these problems are already significantly less significant. And to reiterate, the issue was not the exploitation of programs but exploitation of the virtual machines isolation mechanisms.
MLC uses Apache TVM to generate and autotune the webGPU code, and its respectably performant.
Remember when people would boast about how little lines of code their program had or how little memory their programs used?
LLMs are especially tricky for WebGPU because the good models are so RAM/VRAM heavy.
Already done.
So what I know is that this generates images via browser rather than server. The only thing I can think of is not having to refresh the page in order to change an image or generate a new image. Which... hmm, well, that could mean websites whose visual design changes in real-time? And maybe changes in a way that would be functionally relevant/useful? That does seem pretty cool, although I'm not sure how useful Stable Diffusion is for generating UI components/visual aspects of a site.
Sensitive prompts will not leak to some remote party.
That's a fascinating possibility, but we're very far from that world right now: elsewhere in the thread it's mentioned that this actively uses 8 GB of RAM. And I doubt many web designers would accept the risk that a model misinterprets a prompt, produces distorted output (like the wrong number of fingers on someone's hands), or accidentally produces sexual or violent content in a context where it's not intended.
For many generative image models today, people often pick the best of a dozen or more images, and the others that they throw away may actually be quite bad.
The quality and predictability of the models would need to be significantly higher than it is now in order to routinely dynamically illustrate web sites.
But I don't want to say that we'll never get there. All of the recent models are doing things now that would have been considered inconceivable just a few years ago. (Compare https://xkcd.com/1425/ where it may even be a challenge to explain the issue behind the joke to some younger readers!)
Perhaps we could have an ad-free internet after all.
A miner would probably just generate enough work to pay the engineers and infrastructure.
(Not going to install Chrome for a random shiny thing)
What causes this?
I don't know the specifics of why the slowdown is so extreme in this case, usually it has a negligible impact. But I'm guessing it's related to what I wrote above.
This isn't unique to the web, either, adding the verbose flag to most linux file utilities and then operating on a large set of files will be slower than without the verbose flag, too, just because printing to stdout takes time.