MLC-LLM: GPT/Llama on consumer-class GPUs and phones
github.com
github.com
"The marvel is not that the bear dances well, but that the bear dances at all."
The model we are using is a quantized Vicuna-7b, which I believe is one of the best open-sourced models. Hallucination is a problem to all LLMs, but I believe research on model side would gradually alleviate this problem :-)
Those two events are causally related. The OS has to throttle down the CPU or else it will overheat and malfunction.
It is one of the reasons why heavy number crunching is often performed on the cloud instead.
Citation needed :)) There are economies of scale and various optimizations that are just not possible with dedicated machines.
Anecdotal evidence: a company I worked in had a dedicated DC with hundreds/thousands of machines that mostly ran SQL queries on petabytes of data (any query would take ~5-30 minutes). Eye-watering budget and whole teams to maintain the cluster... They switched to GCP/BigQuery, got queries that ran in seconds at a fraction of the budget.
Economies of scale go in the other direction usually.
Cloud hosting only makes sense in a case that your capacity needs are changing rapidly and unpredictably, or the case where you’re so big that the cloud hosting company is effectively a department. Any other time it will almost certainly be cheaper to self host.
Nowadays many cell phone application processors also have dedicated hardware to accelerate neural nets, but they will always be limited by thermal constraints.
If you're going to be that loose with your definition of hallucination, I'd really hope you apply it to yourself. That's a level of introspection that you need to avoid letting biases that all of us have deep down end up affecting your higher order function.
We humans rely on a ball of biases, interpolations and extrapolations to function, LLMs don't have a monopoly on that.
Our approach leverages TVM Unity, a machine learning compiler that supports compiling GPT/Llama models to a diverse set of targets, including Metal, Vulkan, CUDA, ROCm, and more. Particularly, we've found Vulkan great because it's readily supported by a wide range of GPUs, including AMD and Intel's.
BTW, an interesting data point from Reddit that it also works on steam deck: https://www.reddit.com/r/LocalLLaMA/comments/132igcy/comment....
I have 64gb RAM (not gpu just normal), I’d like to see proof of concepts that the bigger models can be fine tuned and have far more accepted results, or to know if we’re completely going the wrong direction with this
If I had the gumption (and a data set) I could afford to spend a few hundred bucks to fine tune a model for shits and giggles and I’m just a Random Internet Dude.
I’m all for it, Any Day Now™ I have this idea I want to try and having these people do all this optimization work will probably make it affordable to attempt given I don’t actually know what I’m doing so there will be a whole lot of “yeah, that doesn’t work” going on.
Long term maybe Nvidia is able to release release some integrated ARM chip, but I’m not holding my breath.
The Risc V area, now there we can talk about disruption long term
If I'm wrong I'd love a pointer to the docs about it!
It should also be noted that the method of quantization makes a big difference. In particular, if you were experimenting with llama.cpp, their original take on it was considerably inferior to GPTQ. And for the latter, parameters such as group size can also make a difference.
But for some reason it dramatically slows down after a few messages
Edit:
Oh no, this one also gives lectures instead of answering questions.
https://i.imgur.com/eiuGzK4.jpg
I'm afraid, in near future the only organic content on the internet would be only the type of content that LLMs refuse to generate.
I've been tempted to try it myself, but then the thought of faster LLaMA / Alpaca / Vicuna 7B when I already have cheap gpt-turbo-3.5 access (a better model in most ways) was never compelling enough to justify wading into weird semi-documented hardware.
There are already efforts underway with GPTQ libraries but I have found they incur a substantial performance penalty, with the benefit of consuming much lower VRAM.
EDIT: I had a look at the repo, it appears the Vicuna model is using 3bit quantization.
It's on our plan, but we haven't looked to enable them by default in the first place, mainly because we wanted to demonstrate it running on all GPUs including old models that don't come with TensorCore at all.
Because LLMs are expensive to host. It's more of a case that no one wants to run these on their PCs and it's the cheapest if it ends up running on client PCs instead of your own PCs. Not all use cases of LLMs need a super powerful model that is always up to date.
You get to decide what is appropriate or not.
It works offline.
It can be used to by applications without the having to use an external service.
This can be important for a number of applications (I am thinking about open source games and the modding community right now, but it is just an example)
1) for people who want to make money using AI but can’t afford to pay for a LLM or servers, they can push that cost to the end user.
2) for people who want to generate porn or spam (probably, also to make money)
The privacy thing is complete nonsense. If you want a private server, rent your own private server. If you’re worried AWS is spying on you, you’re paranoid.
This is about money, making money and being cheap, not about good will.
So, you’re right; from a consumer perspective it’s pretty meaningless.
More to the point, I find it absolutely bizarre that one couldn't come up with quite a few reasons to have this be more private, whether personal or business.
Your personal or business stuff in other people's hands is generally not optimal or preferable, especially when more private options exist.
Are these the only supported models as of now? https://github.com/mlc-ai/mlc-llm/blob/d3e7f16c54238b7da5e78...
The way we make this happen is via compiling to native graphics APIs, particularly Vulkan/Metal/CUDA, making it possible to run with good performance.
I've scoured the web page for ram requirements for the various models but I can't see anything, will it be able to run let's say the 30B open assistant llama or 65B raw llama model on a consumer gpu (let's say 3060 with 12gb vram) using this?
Not trying to take anything away, but the readme etc is very lacking in actual technical details I feel without reading through the code or actually testing it.
We are expanding the coverage to more models, particularly, Dolly and StableLM are just around the corner, needing some clean up work.
As a fresh new project, right now we are starting to collect data points of which GPU models are supported well and fixing issues being reported. Please don't hesitate to report in our github issue!
In any case I am happy to see these projects taking form. Perhaps one can eventually make the level of quantization dynamic based on the available vram etc :)
I will definitively play around with it (on linux though, not a phone!)
It runs a lot faster if you compile with cuBLAS (nvidia) or clblast (other). GPU vram doesn't matter much since it doesn't offload the model to vram.
https://github.com/mlc-ai/binary-mlc-llm-libs
Is the code from which these are built available somewhere? How does one go about building one for their own model?
What are you using local LLM's for?
So far, I've been only able to come up with:
- Aid in coding (which always ends up in chatGPT)
- Summarizing short articles
- whisper-ai + langchain + ffmpeg allows for some great video summarization (especially with non-english LORA's for us non-natives)
- generating stable diffusion prompts
Also, you hint at those many ideas, could you elaborate on that a bit? I'll be playing with LLMs in near future, might as well do something useful with them
without getting into too much detail my job supports business to people interactions. my use case is training the LLM to assist the business agents. if it can give real time information that's helpful to the agent while causally listening to the conversation thats a pretty big game changer. also i want to use it for staffing decisions since it can view historic data and make recommendations for the future.
Asking for a friend…
> ASSISTANT: Understood! I'll be here to answer any questions you may have in the shell terminal. Let's get started!
> USER: ls
> ASSISTANT: I'm sorry, I can't execute the command you entered as it is a shell command which I am unable to execute as a terminal.
I think it needs a bit more work
Otherwise you should've asked it to pretend to be a terminal and generate fake command output for various common unix binaries.
It's like someone builds a CPU with a floating point unit specifically aimed at CAD software. Then someone else comes and builds a floating point unit for physics simulation. Then someone else ...
Can't we just get a generic compute model, and make that work everywhere? And don't we already have that, e.g. CUDA?
The history of GPGPU in a nutshell…
There are a few “generic compute models” but no incentive for the GPU manufacturers to support them over their proprietary model. Everyone could natively support Cuda and Vulcan and Metal and OpenCL and SPIR-V and…think I’m forgetting one but you get the point.
(on the other hand I wouldn't be surprised if they didn't come with it neither, due to the difficulties of it)
+
Is there any optimization for LLM to run on RTX cards? 40XX,30XX I found out tha LLAMA.CPP is nice but I want to take advantage of my graphic cards also, and didn't found any documentations...
I've rented a server but it has no GPU. Does MLC work well through only CPU inference?
I'd like to get it set-up with langchain if it does work well