GPU-Accelerated LLM on an Orange Pi
blog.mlc.ai
blog.mlc.ai
A bit like you can self host your apps at home on a Umbrel box.
I wonder if the NVIDIA Jetson serie would be the hardware that makes the most sense?
I struggle to see any sizable market for the tinybox, but I wish them good luck.
The Vision is a whole different universe… with their cash and position they could (and may) take 10 cracks at it.
It doesn't take much testing to come up with the ideal power limit for your given workload(s).
[0] - https://www.pugetsystems.com/labs/hpc/nvidia-gpu-power-limit...
Smartphones aside, little Ryzen 6000 boxes would be OK.
Used DDR5 laptops with a little discrete GPU would be even better. I have one with a broken screen that may be dedicated to this very task.
You could maybe run something on the 4GB Jetson Nano?
But very slow: https://www.reddit.com/r/LocalLLaMA/comments/12c7w15/the_poi...
A 32GB+ ddr5 laptop with a dGPU and some RAM will (IIRC, just barely) do llama 70B for far less money and a similar TDP.
They have their "special place" for certain applications but the software is a mess (old driver and CUDA versions, Jetpack is still based on Ubuntu 20.04), the ARM cores are (very) weak relative to ARM flagships, and the performance/price ratio makes no sense unless you really need the form factor and energy efficiency. Oh yeah and the SD card storage is typically frustratingly slow and often unreliable.
A $500 Jetson Nano devkit with 8GB of shared RAM has roughly 10% of the performance of even ancient cards like the GTX 1070 (8GB VRAM alone) that you can throw in a random used x86_64 tower or whatever for $300 all-in. For the extra couple of hundred dollars difference you can get a more recent GPU with higher compute capability, extra storage, more system RAM, whatever. Significantly higher power usage and larger form factor but with power optimization, scaling, etc this approach makes very little difference in practice for occasional inference workloads for "typical home use". I have this configuration idling (with models loaded) at 20 watts, scaling up to around 100-150w for however long given inference loads take to execute, and then scaling back down.
You can kinda see some of the supported backends gated behind flags in the cmake file: https://github.com/mlc-ai/relax/blob/mlc/CMakeLists.txt
Unbatched token generation is basically RAM bandwidth limited, as the entire model has to be cycled through for each token. I bet theoretical performance is similar to the GPU, albeit with much lower power consumption.
How many users would realistically be able to use it at the same time when running on such a device? I am interested in its scalability.
- memory bandwidth
- interconnect bandwidth between the CPU and GPU
- interconnect bandwidth between GPUs
- thermals and power if you're doing a good job of optimizing the rest
I don't see how a batching mechanism would improve on any of those, superficially it looks as though that would make matters worse rather than better. Can you explain where the advantage comes from?
https://groq.com/wp-content/uploads/2020/05/GROQP002_V2.2.pd... the "batching" section of https://docs.nvidia.com/deeplearning/tensorrt/archives/tenso... https://le.qun.ch/en/blog/2023/05/13/transformer-batching/
If you're doing inference on multiple prompts at the same time by doing batching, you don't take more time in streaming. But each streamed weights gets used for, say, 32 calculations instead of 1, making better use of the GPU's compute resources.
You'll get a log_2-based scaling efficiency with nearly any batchsize increase, pending some limitations (memory, etc).
That should be enough at least to roughly sketch it out.
Otherwise ~1.5 tokens/s is definitely the minimum you'd want streaming tokens to a single person.
You could run it using OpenVINO on IntelCPUs, but the performance would probably take a hit. It would be a lot easier though since you can just use ggml.
I'm crediting llama.cpp of all things for being the boost to really up the ante on open source model compilation.
Whatever it is, at least, many of these open source things feel like they just 'happen', as an eventuality, but in order for that to happen it takes a lot of work from a lot of people! Really happy to see the dream of this particular kind of democratization opening widely! :)
Also, it doesn't seem to say anything about the input model's format? pytorch weights? onnx?
The input is a model (weights + runtime lib) compiled via the mlc-llm project: https://mlc.ai/mlc-llm/docs/compilation/compile_models.html
However it seems that the Mali-G610 MC4 is in about the same range as the cheaper models of the old Jetson Xavier.
The newer Jetson Orin models have much faster Ampere GPUs with between 1024 and 2048 FP32 ALUs. Nevertheless, the various Jetson Orin models have a price between 4 times and 13 times higher than a SBC with RK3588 and 16 GB DRAM (especially all the Orin models with more than 8 GB DRAM are very expensive) and the ratio between their prices is much greater than the ratio between their performances.
Any small computer with AMD Phoenix offers a much better GPU performance per dollar than any NVIDIA Orin. The use of NVIDIA Orin is justified only when one needs a device that is qualified for an automotive environment.
> Any small computer with AMD Phoenix offers a much better GPU performance per dollar than any NVIDIA Orin.
> The use of NVIDIA Orin is justified only when one needs a device that is qualified for an automotive environment.
Nvidia Orin will use significantly less energy though.
Not really.
A Ryzen 7 7840U has a GPU with 768 FP32 ALUs @ 2.7 GHz and a NPU that can do 10 TOPS and it has a default TDP of 28 W.
The top Jetson AGX Orin models consume up to 60 W or 75 W, but they are so expensive that it does not make sense to compare them with a computer with 7840U and 32 GB of LPDDR5x-7500 that costs 3 times less.
A comparison that makes more sense is with a Jetson Orin NX 16GB (still significantly more expensive), which has a GPU with 1024 FP32 ALUs @ 0.918 GHz and it has a default TDP of 25 W.
For graphics tasks, Jetson Orin NX would be several times slower than an AMD Phoenix, due to its low GPU clock frequency and much slower CPU cores. The same is true for any programs executed on the CPU cores.
On the other hand, for AI inference, Jetson Orin has very fast tensor cores, so it can be many times faster than an AMD GPU or an ARM GPU, i.e. Jetson Orin NX 16 GB is claimed to be able to do 100 TOPS, so if this is the main intended application it can be worthwhile. Nevertheless, the usefulness of the Jetson Orin models for AI inference is diminished by the fact that their price increases very steeply when more memory is desired.
It really doesn't. It doesn't even know what it knows and what it doesn't know. Without ways to check up on whether what it told you is true or not you may well end up in more trouble than where you were before.
It’s less likely to hallucinate this way.
The Wikipedia patch doesn’t make much sense to me.
What percent of the important questions being asked in this doomsday scenario actually have their answer in Wikipedia?
If 50% of the time you are left trusting raw LLaMa, then you don’t really have a decent solution.
I do appreciate the sentiment tho that future or finetuned LLMS might fit on an RPi or whatever, and be good enough.
I'm not sure if this can work or not but it would be nice to see a trial, you could probably do this by hand if you wanted to by breaking up the answer from an LLM into factoids and then to check each of those individually, and to assign a score to them based on the amount of supporting evidence for the factoid. I'd love that as a plug-in to a browser too.
That was my personal experience in general with ChatGPT as well as LLaMa1/2.
They trained up their own LLM, but from the text it seems like it might be possible to use any LLaMA-style LM without retraining. Not sure though, need to give it a proper look.
There exists (at least) a project to train and query an LLM on local documents: privateGPT - https://github.com/imartinez/privateGPT
It should provide links to the the source with the relevant content, to check the exact text:
> You'll need to wait 20-30 seconds (depending on your machine) while the LLM model consumes the prompt and prepares the answer. Once done, it will print the answer and the 4 sources it used as context from your documents
You will have noticed, in that first sentence, that it may not be practical, especially on an Orange Pi.
This, without taking into account reasoning and consistence. And already this notion that I picked randomly is not without issues: dollars how computed? And, it is not difficult to state that Columbus reached the Caribbeans in 1492; more complex to "decide" the year of the siege of Troy out of the many dates proposed.
But already at the simplified level of determined clear notions: if LLMs are told that "A is B", and in absence of inconsistency in the training corpus, what is the failure rate (i.e. then outputting something critically different)?
> ways to check up
Some LLMs work as search engines, outputting not just their tentative answer but linked references. A reasonably safe practice at this stage is to use LLMs that way: ask then use the output to check the reference.
English Wikipedia will fit on an SD card. It's more valuable and more practical.