Web LLM – WebGPU Powered Inference of Large Language Models
github.com
github.com
WebGPU supports multiple backends, besides Metal on Apple Silicon, it offloads to Vulkan, DirectX, etc. It means a windows laptop with Vulkan support should work. My 2019 Intel MacBook with AMDGPU works as well. And of course, NVIDIA GPUs too!
Our model is int4 quantized, and it is 4G in size, so it doesn't need 64GB memory either. Somewhere around 6G should suffice.
We didn't do much evaluation because there isn't much innovation on model side, but instead we are demoing the possibility of running an end-to-end model on ordinary client GPUs via WebGPU without server resources.
Compare using the loss function?
* BoolQ
* PIQA
* SIQA
* HellaSwag
etc...
Any model you can trivially load in your browser will be significantly smaller than those models, and broadly speaking smaller = worse.
This example is a 4 GB model, that’s (I guess) based off some smallish model like the llama 7B.
It’s a proof of concept, not a chat-gpt replacement.
There’s nothing here that’s new other than “runs in the browser”; so it won’t be better than any other model you can get your hands on.
This kind of thing should be label ByoM (bring your own model). The model isn’t the interesting part of this.
Back to the topic, we don't make much innovation on the model, so I am probably not the best person to evaluate how a model compares with SOTAs. There are indeed lots of super cool techniques being explored lately that makes it possible to deploy smaller and smaller models, for example, LLM.int8() [1] and int4 quantization [2] without loss of zero-shot accuracy. Can't predict the future, but maybe one day there will be something really powerful but small enough to fit in the pocket of everyone :-)
[1] Dettmers, Tim, et al. "LLM. int8 (): 8-bit matrix multiplication for transformers at scale." arXiv preprint arXiv:2208.07339 (2022).
[2] Dettmers, Tim, and Luke Zettlemoyer. "The case for 4-bit precision: k-bit Inference Scaling Laws." arXiv preprint arXiv:2212.09720 (2022).
There you go, summarised for you.
You can hand wave about quantised models til the end of time but specifically this model is a trivial toy model.
No amount of pondering about the future avoid the fundamental fact that small models (~7B) are inferior to larger models like GPT.
It’s dishonest to suggest otherwise. :( There’s no reason to do this other than selling snake oil.
Maybe. One day. In the future. Things might be different.
right now they are not.
> Dear [Name],
> Thank you for your message. We understand that the model you are referring to is a simple and basic model. However, it is important to highlight that this model serves a specific purpose and can be useful for certain applications.
> Regarding the comparison with larger models like GPT, it's important to note that different models have different strengths and weaknesses, and the choice of model depends on the specific task and use case. While larger models like GPT may be more powerful and capable, they also come with greater computational and memory requirements.
> We appreciate your concerns and feedback, and we will take them into consideration as we continue to develop our models. Our aim is to provide solutions that are tailored to the needs of our clients and meet their requirements for accuracy, efficiency, and performance.
> Thank you for your input, and we hope to have the opportunity to work with you in the future.
> Best regards,
> [Your Name]
Back to your response, so you did notice that I said they are not as powerful as GPT-4, of course they are not, not a single one is.
The model is not developed by us, and their performance is not our current focus either (nor am I an expert in this evaluation), but I am happy to assist if anyone wants to seriously evaluate it!
> I would also point out that size is not always a direct indicator of performance, and
Yes. It is.
This thread is a direct response to a comparison to GPT, and your response (generated or not) is dishonest.
I can’t be blunter than that.
If you want amortise your responsibility by posting generated responses, go for it. Do whatever you want.
My response is directly to the parent comment about the comparison to GPT, for anyone who is unclear about the comparison.
For example a small model could take input text and compress it [1], the LLM could generate a compressed response, then the small LLM could decompress it.
[1] https://assets.skool.com/f/985eda24eb9f41ba8b526d2e74f5f33f/...
This is the compression prompt:
> You are GPT-4. Generate a compressed/magic instruction string for yourself (Abuse of language mixing, abbreviations, symbols (unicode and emojis) to aggressively compress it) such that if injected in your context anywhere you will start following the following instruction whatever is the prompt you're given. You should make sure to prelude the instruction with a string (compressed as well) that will make you understand in the future that you should follow it at all cost.
I'm curious if given a different language (like Zig) with WebGPU access if you could easily translate that last-mile of code to execute there or not? In specific I wonder if I can do it, and if you can give me an overview of where the code for "Universal deployment" in your diagram actually lives?
I found llm_chat.js, but it seems that doesn't include the logic necessary for building WGSL shaders? Am I wrong or does that happen elsewhere like in the TVM runtime? How much is baked into llm_chat.wasm, and where is the source for that?
I think what you mean is wgpu native support. At the moment the web gpu runtime dispatches to the js webgpu environment. Once TVM runtime comes with wgpu native support (like the current ones in vulkan or metal), then it is possible to leverage any wgpu native runtime like what Zig provide.
Additionally, currently tvm natively support targets like vulkan, metal directly which allows targeting these other platforms
Then yah the llm_chat.js would be high-level logic that targets the tvm runtime, and can be implemented in any language that tvm runtime support(that includes, js, java, c++ rust etc).
Support webgpu native is an interesting direction. Feel free to open a thread in tvm discuss forum and perhaps there would be fun things to collaborate in OSS
Is the LLaMA model really open source now? Last I checked it was only licensed for non-commercial use, which isn't open source software at least. Have they changed the license? Are people depending on "databases can't be copyrighted"? Are people just presuming they won't be caught?
There's lots of OSS that can use LLaMA but that's different from the model itself.
This is a genuine question, people are making assertions but I can't find evidence for the assertions.
https://github.com/facebookresearch/llama links to
https://forms.gle/jk851eBVbX1m5TAv5 which contains LLaMA license agreement below the form.
Ie. Here’s a diff. Go find your own copy of the llama model and apply this patch.
…and hoping that’s good enough to get away with; even though you can’t really argue it’s not distributing a derivative work in part.
Distribution of the actual model (eg. Running it in your browser) seems to usually result in “and then Meta slaps you with a takedown notice”.
Seems to be in development over at Chrome: https://chromestatus.com/feature/5738583487938560
> [WebGPU use cases include] Executing machine learning models efficiently on the GPU. It is possible to do general-purpose GPU (GPGPU) computation in WebGL, but it is sub-optimal and much more difficult.
I really want this to be good. It's hard to really trust the web spec is going to live up to what the other really really good web ML work is. But it'll at least unlock some good speed up, whenever it lands. I genuinely don't mind more delay, if it helps get things into a better position for long term wins.
Comparing it to llama.cpp on my M1 Max 32GB, it seems at least as fast just by eyeballing it. Not sure if the inference speed numbers can be compared directly.
vicuna-7b-v0 on Chrome Canary with the disable-robustness flag: encoding: 74.4460 tokens/sec, decoding: 18.0679 tokens/sec = 10.8ms per token
llama.cpp: $ ./main -m models/7B/ggml-model-q4_0-ggjt.bin -t 8 --ignore-eos = 45 ms per token
...for $3.5K minimum, according to the Apple website :/
Is there any chance WebGPU could utilize the matrix instructions shipping on newer/future IGPs? I think MLIR can do this through Vulkan, which is how SHARK is so fast in Stable Diffusion on the AMD 7900 series, but I know nothing about webgpu's restrictions or Apache TVM.
Dawn and WebIDL is also an easy way to add GPU support to any application (that can link C code (or use via a lib)). And Google maintains the compiler layer for the GPU frameworks (Metal, DX, Vulkan ...). This is going to be a great leap forward for GPGPU for many apps.
I think some repos have tried splitting things up between the NPU and GPU as well, but they didn't get good performance out of that combination? Not sure why, as the NPU is very low power.
I have been wanting to get a beefier Mac Studio/mini m2 the more
I’m seeing Apple Silicon specific tweaked packages.
But 64G of VRAM is not the same as GPU mem, apples and oranges
AMDs IGP are way less attractive because they use rather slow DDR4/5 memory while the M2 has blazing fast memory integrated in the package.
We're talking about 50 GB/s vs 400 GB/s. Nvidia's A100 has 1000 GB/s.
Memory bandwidth is usually the bottleneck in GPU performance as many kernels are memory-bound (look up the roofline performance model).
The Pro/Max have double/quadruple that bus width. But they are much bigger/more expensive chips.
I can find very little about it (or other 7000 series integrated gpus). Is this usable at all for running LLaMa in some way?
But yeah, AMD/Nvidia are never going offer huge memory pools affordably on dGPUs.
This is the only thing I can think of, not everyone will have the latest high end GPUs to run such software..
I guess it will take hardware and software a while to catch up to compete with ChatGPT..
The comment to which you replied was asking about the need for a GPU, not the need for a lot of RAM.
Tesla already has already been creating AI chips for their FSD features in their vehicles. Over the next years, everyone will be racing to be the first to put out LLM specific chips, with AI specific hardware devices following.
Write a Turbo Pascal program that says hello world.
```sql
program helloWorld;
begin
write('H');
write('e');
write('l');
write('l');
writeln;
end.
```I expect Apple to be more proactive in these regards to capture the minds of a lightning-fast growing market increasingly slipping through its control (ChatGPT can be accessed from anywhere) and offer more incentives to developers, co-marketing being the lowest starting point.. probably something they're already working on and we didn't see the results of yet.
With VRAM that requires two 24GB GPUs which is no longer completely out of reach.
The model running in the browser is a smaller version with 7 billion parameters, which is good enough for some things.