Or just a powerful apple silicon machine? I've tried dolphin mixtral 4bit on a 36gb ram MacBook m3, and inference is super fast.
Or just a powerful apple silicon machine? I've tried dolphin mixtral 4bit on a 36gb ram MacBook m3, and inference is super fast.
EDIT: I cannot, I need to install ROCm to compile with it, and then install something called hipBLAS, and who knows what else.
EDIT: Oh, no, I have an nVidia GPU, AMD CPU.
Also I only tried it on linux, AFAIK windows is a lot more difficult to get running, if it works at all...
With llama-cpp, I successfully tried various LLMs(e.g. LLAMA 13B, Mixtral etc) with very solid performance. Even for models that don't fit in VRAM completely, performance can be surprisingly solid, as long as you compile with AVX extensions. (and your CPU supports those)
Stable Diffusion via ComfyUI also works very well. However, be aware of VRAM limitations with the larger SDXL variants, especially when running a heavy desktop environment.
Regarding setup guides/links, there isn't a good centralized resource sadly, so some tinkering is needed. Unlike some of those CUDA 1-click solutions, ROCm requires more manual setup, especially for the models only unofficially supported.
Here are a couple of links that might be helpful:
https://old.reddit.com/r/LocalLLaMA/comments/18ourt4/my_setu...
https://old.reddit.com/r/StableDiffusion/comments/ww436j/how...
In general the r/localllama & r/StableDiffusion subreddits are good places to search for info.
Nvidia has significant dominance in the AI space because of their work on software and the overall platform.
With the Jetson line being the sole exception. Use it for what it’s for - a targeted build for an embedded/specific application requiring small size and low power.
The software is a mess. Support for Jetson (generally) is a far afterthought or not considered at all around projects at Nvidia and the broader ecosystem. When it is supported at all it lags behind significantly, using ancient distros (Jetpack), etc. To make matters worse the user base is so (relatively) tiny there are bugs and strange behavior everywhere.
Just don’t do it.
I'd suggest checking what examples are available, see what community is doing, see if what you need had already been tried - https://www.jetson-ai-lab.com
From what I've seen, mainstream LLM libraries like VLLM, llamacpp that use CUDA under the hood tend to work out-of-the-box. And there are tutorials available: https://www.jetson-ai-lab.com/tutorial_text-generation.html. I think that TensorFlow/Pytorch are also well maintained, although I've not checked recently.
Nvidia more broadly has very impressive support for their GPUs. If you look at the support lifecycles for their Jetson hardware over time it's significantly worse. I encourage you to look at what support lifecycles have looked like, with the most "egregious" example being dropping of support for the Jetson Nano in from what I recall was within a couple of years.
Another consideration - Jetson is optimized for power efficiency/form-factor and on a per $ basis CUDA performance is terrible. The power efficiency and form-factor come at significant cost. See this discussion from one of my projects[0]. I evaluated the use of WIS on an Orin Nano that I have and it was nearly 10x slower than a GTX 1070 which is seven years old and is still supported by the latest drivers and CUDA 12 on whatever OS you want.
Nvidia knows what they're doing in terms of productization and the Jetson line should not be seen as some kind of secret hack/unlock for getting CUDA performance with gobs of RAM. In the case of LLMs I wouldn't be surprised at all if CPU beats it and at that point pickup 256GB of RAM or whatever for equivalent cost.
In the end what do I care what people use, I'm offering the perspective and experience of someone who has actually used the Jetson line for many years and frequently struggled with all of these issues and more.
[0] - https://github.com/toverainc/willow-inference-server/discuss...
About a year back, I took that very early version of AGX Xavier, that got released years ago. It wasn't even the version that was officially released. Yet I was able to refresh it to newer Ubuntu without any issues.
Wheels are often not pre-built for aarch64, yes. If you want to compile directly on Nano, disk performance is very important. Sometimes you get I/O bound.
Orin Nano being that slow in [0], it looks like you've been trying it in Aug 2023. It maybe worth re-evaluating on the latest Jetpack, it had transitioned to CUDA 12.2, TensorRT 8.6, cuDNN 8.9. I would expect that recent popularity of ASR/TTS pipelines and LLMs was not completely missed by Jetpack maintainers (there are some tutorials here - https://www.jetson-ai-lab.com ). And recently released JetPack could be optimized a lot more for these workflows.
And your project is very cool! I'd suggest sharing it and your performance numbers (!) with the maintainers of: https://developer.nvidia.com/embedded/community/jetson-proje...
Nice! I'm sorry if I seemed dismissive or even disrespectful, in my experience Jetson certainly has it's place (why I've been using them for years) but compared to "bring your distro, apt-get/.run Nvidia driver" they can be a serious shock for casual users. Then they see the performance...
> Orin Nano being that slow in [0], it looks like you've been trying it in Aug 2023. It maybe worth re-evaluating on the latest Jetpack, it had transitioned to CUDA 12.2, TensorRT 8.6, cuDNN 8.9. I would expect that recent popularity of ASR/TTS pipelines and LLMs was not completely missed by Jetpack maintainers (there are some tutorials here - https://www.jetson-ai-lab.com ). And recently released JetPack could be optimized a lot more for these workflows.
Interestingly WIS was recently bumped to CUDA 12.2, etc and the performance improvements were very marginal. WIS uses Ctranslate2 under the hood (same as faster-whisper) which offers among the best Whisper performance overall but doesn't benefit much from changes in these underlying libraries. In the end even if it somehow magically doubled performance (it doesn't and won't) that still places the latest generation ~$600 Jetson board 5x slower than an ancient yet still fully officially supported ~$100 GPU. Power and form-factor is an issue but for the voice assistant use case a Jetson board barely doing realtime with Whisper medium is unacceptable to me and the vast majority of our users. Our goal is sub one second voice command sessions from end of speech, to command execution, to TTS response and Jetson just can't provide that at any cost.
I'm glad there are community resources for Jetson platforms (which I'm aware of) but their existence underscores my point - you'll notice when perusing through there are often various hoops to jump through whereas anything else is basically "install driver, container toolkit, docker run" and it just works and works performantly. Basically CUDA x86_64 and discrete GPUs is native/expected/developed for, Jetson is almost always a bit of an edge case with rough edges (relatively) all over the place.
> And your project is very cool!
Thanks! In terms of your suggestion I certainly might but in the meantime, overall (based on my Jetson experience) as I said in that discussion I'm very reluctant to officially support the Jetson line with WIS. I'm almost certain it will blow-back on the project and cause support headaches for us while all the while providing a sub-optimal user experience.
I've gotten my Jetson to work well using Yocto to build my own linux distros with correct updated dependencies, libraries and updated jetpack, but it's not for the faint of heart, and that's a whole other ball of yarn. It also takes a few hours to generate a new build every time I need to update some dependency that depends on other dependencies (Yocto maintenance is a full time job in many embedded development shops - you're basically authoring your own distribution).
Treat these devices as what they are: embedded target boards for fixed industrial development (for example, to go into a robot or a car - once that design is finished, Nvidia will expect you to NEVER update any part of the system with an embedded jetson or orin system for years, until you replace the whole thing with their newest model that you buy off the shelf again).
This is standard fare in embedded and robotics space. Do not use these boards for any kind of rapidly moving software development, because it's the wrong tool for the job.
Software for Jetson boards should be viewed as firmware for these embedded/industrial devices. They get installed in a robot, MRI machine, etc with a specific bespoke application targeting what they came with and are never touched again -or- supported by some large commercial firm with the skillsets you describe.
I was as firm/absolute in my original reply as I was because anyone who thinks life with Jetson is similar to life with a discrete Nvidia GPU on x86_64 will be in for a huge shock and 95% of the time it will end up on their shelf in a year or two.
It's one thing when it's the latest random ARM SBC you bought for $50 with no vendor support, it's another thing entirely when you're spending > $600 (or $2000 as this started!!!!) on a Jetson.
[1] https://www.hackster.io/pjdecarlo/llama-2-llms-w-nvidia-jets...
Maybe I will try to get a 32 GB+ LLM running one of those days.
the difference is the bandwidth and the computational power of the APU. M1 Max is roughly similar to a PS5 in terms of overall system design (shader configuration and bandwidth) plus has dedicated AI inference units already (which won't be added to consoles until PS5 Pro launches with RDNA 3.5). It is far more bandwidth than you can get out of a socketed-memory laptop system.
https://twitter.com/Locuza_/status/1450271726827413508
To support that level of performance in a socketed-memory system you will need an extra layer of caching added to the processor to supplement the bandwidth - and maybe still need to go to quad-channel. Those products are Strix and Strix Halo and should be hitting the market over the next year or two but the reality is that the M1 Max was an absurdly powerful laptop, far more potent than even the first-gen 5nm laptops for x86 let alone the other junk you could buy in 2020.
This is the problem with the discourse around apple silicon for the last few years: yeah, they're expensive, but even a loaded-out x86 laptop doesn't get you the same capabilities. Even if the x86 is competitive in some particular benchmark on iso-node you are probably spending more power to do it, and the x86 product comes years after the apple product, and still has a much weaker gpu and less bandwidth (which doesn't just matter for GPU, it matters for compiling and JIT too).
It is incredibly silly to look back on the discourse in 2020-2023 around apple silicon, a lot of reviewers made extremely silly claims about how "even 7nm x86 processors were already competitive with apple silicon" and as the ecosystems have matured it is obvious that even 5nm processors are not quite competitive yet. And they dumped on the SPEC tests and Geekbench that measured this properly, in favor of dumb things like cinebench R23 and so on (it's always cinebench used for this dumb shit tbh, CB R13/R15 were hugely misleading at the zen1 launch too). Let alone things like, you know, compiling or JVM/node workloads...)
(similarly, gotta love the vibe a few years ago of: "threadripper vs mac pro" - did you know that a 64C threadripper with 256GB RAM is actually cheaper than a mac pro loaded out with 2TB!? waow, who knew systems with an order of magnitude less capacity would be cheaper!? https://youtu.be/BH291DQRIOg )
Its kinda like GPT 3.5, with no internet access and slightly less reliable responses, but unrestrained, much faster and with a huge (up to 75K on my Nvidia 3090) usable context.
Mixtral is extremely fast though, at least at a batch size of 1.
Nous Capybara is a popular one. I personally use my own merge of many models, and you can look through the constituent models to see if any interest you: https://huggingface.co/brucethemoose/Yi-34B-200K-DARE-megame...
You can't really use llama.cpp for super long context btw, its just too slow and vram inefficient at the moment.
I heard you can simply install ollama app which uses llama.cpp under the hoods, but has a more user friendly experience.