WebLLM: Llama2 in the Browser
webllm.mlc.ai
webllm.mlc.ai
prefill: 0.9654 tokens/sec, decoding: 3.2589 tokens/sec
Honestly amazed that this is even possible. I haven't even run Llama 2 70B on my laptop NOT using a web browser yet.
Its my understanding that under normal circumstances decoding is memory bandwidth bound which prompt processing isn't due to batching. Is there some quirk in your setup?
It's slow but usable!
prefill: 2.1963 tokens/sec, decoding: 3.4708 tokens/sec
210.6742 tokens/sec, decoding: 20.7758 tokens/sec
Can't tell if its really taking advantage of the all the power of the card, though.
If somebody hasn't tried running LLMs yet, here are some lines that do the job in Google Colab or locally.
! git clone https://github.com/ggerganov/llama.cpp.git
! wget "https://huggingface.co/TheBloke/CodeLlama-7B-GGUF/resolve/main/codellama-7b.Q8_0.gguf" -P llama.cpp/models
! cd llama.cpp && make
! ./llama.cpp/main -m ./llama.cpp/models/codellama-7b.Q8_0.gguf --color --ctx_size 2048 -n -1 -ins -b 256 --top_k 10000 --temp 0.2 --repeat_penalty 1.1 -t 8
[0] : https://en.wikipedia.org/wiki/Atwood's_LawWhat are the exclamation points for though? In a *nix shell they'll expand to a command from history - copy-pasters beware!
Shameless plug: https://github.com/jankovicsandras/ml <- here are some minimal Colab / Jupyter notebooks for absolute beginners.
I just find it amazing how little effort it takes to run an LLM nowdays.
I wonder if this sort of behaviour was more nuanced in the initial model, and something like quantisation has degraded the performance?
For instance, unlike kids, at training time an LLM isn't going to ask “It's not very nice for the parents to abandon their children in the forest, is it?”.
I know conservatives are easily triggered by such caveats, but at the same time, they are literally banning books from library ¯\_(ツ)_/¯
Apparently AI companies can't be bothered to filter out the harmful training data so you end up with this warning every time you reference something even remotely controversial. It paints a bleak future if AI companies will keep producing these censoring AIs rather than fix the problem with their input.
Throw it in the trash, its worthless.
The RedPajama one seems to work alright. It still often ends up getting stuck in a loop, though. GTX 1080. prefill: 34.2173 tokens/sec, decoding: 19.8731 tokens/sec
llama-2 is generating pure nonsense for me, just random letters, numbers, and punctuation Vicuna-v1-7b-q4f32_0 is slightly better. None of the fp16 models work on my GPU, I'm guessing it's a hardware limitation.
At least vicuna generates words, but it sure becomes obvious how much these models are just autocorrect. It reads like I'm tapping the word prediction button on my phone's keyboard.
Human: Once upon a time
AI: gro (2 o r (tella, asan) additionaly, as the combination, as the,as they, as the arrived, as they, the, as the as being, the, were, as the, as themselves, the, as their, the as them, the as their, the arrived, the, as the, as themselves, the, as the, as their, the, the as their, as themselves, as the, as their, the, the as their, the, as their, the, as their, the as their, the, as their, the, as their, the, as themselves, the, as their, the, as their, the, as their, the, as their, as their, the, as their, the, as their, the, as their, as their, the, as their, as their, as their, the, as their, the, as their, the, as their, as their, as their, the, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, as their, asprefill: 16.9337 tokens/sec, decoding: 0.4631 tokens/sec on a Radeon 6600XT on the 7b default model.
Does feel something isn't quite right as it's only using a few % of GPU/CPU - though it is using the AMD GPU! Which I have never managed to get working in Windows or Linux with llama.cpp directly.
I wonder if using WebGPU somehow would avoid all the horrendous problems of GPU support in LLMs, as it seems there is some sort of semi-working abstraction layer here which works across M1/M2, Nvidia and AMD?
Generate error, Error: Chat module not yet initialized, did you call chat.reload?