Llama.cpp: Full CUDA GPU Acceleration
github.com
github.com
I find it interesting that it's an example of an ML software that's totally detached from Python ML ecosystem and also popular.
Is Python so annoying to use that when a compelling non-Python solution appears, everyone will love it? Less hassle? Or did it take off for a different reason? Interested in hearing thoughts.
For me it's less about the language and more about the dependencies. When I want to run a Python ML program, I need to go through the hassle of figuring out what the minimum version is, which package manager/distribution I need to use, and what system libraries those dependencies need to function properly. If I want to build a package for my distribution, these problems are dialed up to 11 and make it difficult to integrate (especially when using Nix). On top of that, those dependencies typically hide the juicy details of the program I actually care about.
For something like C or C++? Usually the most complicated part is running `make` or `cmake` with `pkgconfig` somewhere in my path. Maybe install some missing system libraries if necessary.
I just don't want to install yet another hundred copies of dependencies in a virtualenv and just hope it's set up correctly.
This is exactly why I hate Python. They even have a pseudo-package.json style dependencies file that you should supposedly be able to just run "install" with, but it NEVER works. Not once have I ever downloaded someone's Python project from github and tried to install the dependencies and run it has it ever gone smoothly and without issue.
The Python language itself may be great, I don't know, but I'm forever put off from learning or using it because clearly in all the years it's been around they have yet to figure out reproducibility of builds. And it's obviously possible! JavaScript manages to accomplish it just fine with npm and package.json. But for some reason Python and its community cannot figure it out.
I'd really like to see more Docker images (images, not Dockerfiles that fail to build). Maybe flatpack or snap packages do the trick, too.
Same here. If it can't resolve dependencies or whatever, then there will almost certainly be some kind of showstopping runtime error (probably because of API changes or something). I avoid Python programs at all cost nowadays.
`pip` is the package manager that almost works.
Python is the language that almost supports package distribution.
I'll keep using `apt` on vanilla Debian.
I'm not going to install thousands of dodgy-looking packages from pip, the only documentation for which is a Discord channel full of children exchanging 'dank memes'.
I like Python, but I simply do not trust the pip ecosystem at this point (same for npm, etc.).
This made me laugh. It’s true, isn’t it? That’s really what we deal with day to day (for me in the js world, the create react app dependencies make my head spin)
I don't mean to detract from your main point re Python dependencies, but I find this about C to be rarely true. `make` etc build flows usually result in dependency-related compile errors, or may require a certain operating system. I notice this frequently with OSS or academic software that doens't ship binaries.
Which are displayed to me during the `./configure` step before the `make`, and usually require me to type "apt-get install [blah] [blah] [blah]", and to run configure again.
I'm glad make, etc., works for yo. But for me, neither C, C++, nor Python are particularly enjoyable dependency wise.
Even the more complex and annoying scenarios in native-land for dependency management still feels positively idyllic in comparison to Python venvs.
It also helps that even minor versions introduce breaking changes.
I doubt anyone really knows Python that well, unless they are on the core team.
Of course, that worked because 1) I was really only doing one project, not juggling multiple ones, 2) there weren't all that many dependencies (Numeric, plotting, etc.), and 3) I was already up to my eyeballs in the build system with SWIG and linking to the actual compute code, so I knew my way around the system.
But every now and then I just shake my fist at the clouds then mutter darkly about just installing the dang thing and maybe not taking on so many dependencies. :-)
The annoyance with venvs is you have to create and activate them. In contrast for cargo (or stack or dotnet or yarn or pipenv or poetry), you just run the build tool in the project directory.
Another limitation of venv is it doesn't solve the problem of pinning a versions of Python, so you need another tool.
Dependency problem still happens even with C/C++.
With a boring C project, if it compiles it probably works without hassle.
Feels validating that other people have these thoughts too and I'm not just some old fart.
The answer is "yeah use this other software to make it work in an isolated way because the whole ecosystem is actually broken" and that's somehow acceptable
Otherwise you not only bind to system architecture and OS, you also bind yourself to a distribution.
I find that Automatic1111 plugins tend to not share dependencies and instead redownloads them for their own use. Can make your hdd cry because some of these are larger models. Advantages and disadvantages probably...
There are package managers for C and some are quite good. But for most projects you are quite dependent on the package manager of your distro to supply you a fitting foundation. Sometimes it is easy, but if there is a problem, I think handling C is far harder than python. And I write quite a bit of C while I can only perhaps read python code.
No code is completely platform independent, especially a stable diffusion project, but Python is still more flexible as C by a long shot here.
Of course Llama is great. Time to get those LLMs on our devices for our personal dystopian AIs running amok.
I don't understand why things are so complicated in Python+ML world.
Normally, when I have a Python project, I just pick the latest Python version - unless documentation specifically tells me otherwise (like if it's still Python 2 or if 3.11 is not yet supported). If the project maintainer had some sense, it will have a requirements list with exact locked versions, so I run `pip install -r requirements.txt` (if there is a requirements.txt), `pipenv sync` (if there is a Pipfile), or `poetry install` (if there's pyproject.toml). That's three commands to remember, and that's not one just because pip (the one de-facto package manager) has its limitations but community hadn't really decided on the successor. Kinda like `make` vs automake vs `cmake` (vs `bazel` and other less common stuff; same with Python).
External libraries are typically not needed - because they'll be either provided in binary form with wheels (prebuilt for all most common system types), or automatically built during the installation process, assuming that `gcc`, `pkgconfig` and essential headers are available.
Although, I guess, maybe binary wheels aren't covering all those Nvidia driver/CUDA variations? I'm not a ML guy, so I'm sure how this is handled - I've heard there are binary wheels for CUDA libraries, but never used that.
Ideally, there's Nix (and poetry2nix) that could take care of everything, but only a few folks write Flakes for their projects.
> Usually the most complicated part is running `make` or `cmake` with `pkgconfig` somewhere in my path
Getting the correct version of all the dependencies is the trickiest part as there is no universal package managers - so it's all highly OS/distro specific. Some projects vendor their dependencies just to avoid this (and risk getting stuck with awfully out-of-date stuff).
> Maybe install some missing system libraries if necessary.
And hope their ABIs (if they're just dynamically loaded)/headers (if linked with) are still compatible with what the project expects. At least that is my primary frustration when I try to build something and it says it doesn't work anymore with whatever OS provides (mostly, Debian stable fault lol). It is not exactly fun to backport a Debian package (twice so if doing this properly and not handwaving it with checkinstall).
Relevant to "AI, Python, setting up is hard ... nix", there's stuff like:
It's fair to say part of the blame is on Nvidia, but wow is it frustrating when you have to find eclectic mixes.
Many pythonic projects not only in ML world tend to be just developers experiments, so to be run as an experiment, not worth to be packaged as a stable, released program...
Oh, BTW projects like home-assistant fell in the same bucket...
Worse, because Python is a dynamic language, you also have to deal with all of that complexity at deployment time. (Vs C/C++/Zig/Rust where you can just ship the compiled binary).
This wasn't true for decades, `virtualenv` was de-facto standard isolation solution (now baked in as `python -m venv`, still de-facto standard), and `pip` is the package manager (we don't talk about setuptools/distutils, ssh!). If someone still used system-wide packages that was either because a) they were building a container or some single-purpose system; or b) they were sloppy or had no idea what they're doing (most likely, following some crappy tutorial). Or it was distro people creating packages to satisfy dependencies for Python programs - but that's a whole different story (and one's virtualenv shouldn't inherit system packages unless it is really really necessary and iif it makes sense to do so).
The problem started when one needed some external non-Python dependencies. Python had invented binary wheels and they're around for a while (completely solving issues with e.g. PostgreSQL drivers, no one needs to worry about libpq), but I suppose depending on specific versions of kernel drivers and CUDA libraries is a more complex and nuanced subject.
> Vs C/C++/Zig/Rust where you can just ship the compiled binary
Only assuming that you can either statically link, or if all libraries' ABIs are stable (or if you're targeting a very specific ABI, but I've had my share of "version `GLIBC_2.xx' not found"s and not fond of those).
In a similar spirit, any Python project can be distributed as one binary (Python interpreter and a ZIP archive, bundled together) plus a set of zero or more .so files.
Right; but python itself doesn’t check your local virtual environment unless you “activate” it (ugh what). And it can’t handle transitive dependency conflicts, like node and cargo can. Both of those problems stem from python assuming that a simple, flat set of dependencies are passed in from its environment variables.
Or simply because they are packagers for some distro and they user want a simple way to pull-in some software by it's name, while the upstream devs imaging people cloning their public repo and run the software from the checkout in their own home, with regular pull, regularly rebuilding the needed surroundings...
Not to talking about modern systems/distro with not-really-posix vision like NixOS or Guix System...
> In a similar spirit, any Python project can be distributed as one binary
A single 10+Gb binary :-D
If you don't use virtual environments in python, isn't it basically the same in python? Just run `pip install` and maybe install some missing system libraries if necessary. In practice, it's not that simple in either language, and "maybe install some missing dependencies" sweeps a lot of pain under the rug.
Having used both C++ and Python for some time, the idea that managing C++ dependencies is easier than venv and pip install is one of the moments you wonder how credible is HN opinion on anything.
> a compelling non-Python solution appears
Confusing a large ML framework like pytorch that allows you to experiment and develop any type of model with a particular optimized implementation in a low level language suggests people are not even aware of basic workflows in this space.
> also popular
Ofcourse its popular. As in: People are delirious with LLM FOMO but can't fork gazillions to cloud gatekeepers or NVIDIA so anybody who can alleviate that pain is like a deus-ex-machina.
Ofcourse llama.cpp and its creator are great. But the exercise primarily points out that there isn't a unified platfrom to both develop and deploy ML type models in a flexible and hardware agnostic way.
p.s. For julia lovers that usually jump at the "two-language problem" of Python: here is your chance to shine. There is a Llama.jl that wraps Llama.cpp. You want to develop a native one.
This is because most people don't care about developing the project, just using it. So they don't care what the dependencies are, just that things work. C++ might be more difficult to handle dependencies to build things, but few people will look into hacking on the code before checking to see if it's even relevant.
Not sure if that was half sarcastic, but from experience, anything touching c++ & CUDA has the potential to devolve into a nightmare on a developer-friendly platform (hello kernel updates!), or worse, into a hair-pulling experience if you happen to use the malware from Redmond (to compound your misery, throw boost and Qt dependencies in the mix).
Then again, some of the configurations steps required will be the same for Python or C++. And if the C++ developer was kind enough to provide a one-stop compilation step only depending on an up-to-date compiler, it might, indeed, be better than a Python solution with strict package version dependencies.
I think you are generalizing. I do not hate on Python the language but this ML projects are a very , very terrible experience. Maybe you can lame the devs of this ML projects, or the ones of the dependencies but the experience is shit. You can follow a step by step instruction that worked 11 day ago and today is broken.
I had similar issues with Python GTK apps, if the app is old enough then you are crewed because that old GTK version is no longer packaged, if the app is very new then you are screwed again because it depends on latest version of some GTK wrapper/helper.
Its deployment story is where this gap frequently shows. Desktop apps at best passable, whereas e.g. android apps practically non-existing despite the efforts of projects like kivy.
And some jumped on the started bandwagon unknowingly, and never got off of it.
I guess since these models are meant to be run "client side" or "at the edge" or whatever you want to call it, it helps if they can be neutrally used with just about any wrapper. Using them from Javascript instead of Python is sort of huge for moving ML off the server and into the client.
I haven't really dipped my toes into the space until llama and whisper cpp came along because they dropped the barrier extremely low. The documentation is simple, and excellent. The tools that it's enabled on top like text-generation-webui are next level easy.
git clone. make. download model. run.
That's it.
Building (training) machine learning, deep learning models are much more complex , order of magnitude complex than just running the models and doing that in C or C++ would take you years which would take just a few month with python.
And complexity of `pip install` is nothing compared to that.
That's why no real ETL+Deep learning, training work is done in c or c++.
The big shift is Jupyter, but that's mainly for exploratory programming. If you already know what you're doing, there's no reason why C++ should be worse than Python for training. It's likely that most ML engineers do not have experience with C++.
I don't have that experience either but from what I've seen, C++ is very powerful, so once you subtract the jupyter, there's not really too much left.
BTW: You cannot use DOM/CSSOM from C++, the only API is JS, so your argument is theoretical.
Take a look at rust’s Cargo for what a modern package manager should look like. Or deno / Go if you swing that way.
Which old language gets package management right? None of them. None of them get it right.
And sure - conda / venv / CMake / etc help. But last century’s bad design decisions still shine through.
The flip side is like you said... You will just have to reimplement everything yourself and then you can never worry about dependencies again! Just hope you didn't introduce some obscure security issue in your hashmap implementation .
If I have use a dynamic language, at least make it battery included like Ruby. Sure it's also not performant but I get something back in exchange.
Python sits in a very uncomfortable spot which I don't find a use for. Too verbose for a dynamic language and not performant enough compared to a more static language.
The testing culture is also pretty poor in my opinion, packages rarely have proper tests, (and especially in the ML field)
1) function decorators etc have made the code unreadable
2) while code is succinct, a lot of abstraction is hidden in some C/C++ language binding somewhere, so, when there is a problem, it is hard to debug
3) Pytorch has become a monolithic monster with practically no-one understanding its functionality e-2-e
For everything else, I rather use compiled languages with JIT/AOT toolchains, and since most "Python libraries" for machine learning are actually C and C++, any language goes, there is nothing special about Python there.
Its always been obvious that ML’s marriage to python has always been credential ladened proponents in tangentially related fields following group think.
As soon as we got a reason to ignore those PhDs, their gatekept moat evaporated overnight and the community of [co-]dependencies became irrelevant overnight.
It didn’t make python obsolete then and it won’t now.
Training (orchestration, architecture definition etc.) and data munging (scraping cleaning analyzing etc.) are much easier with python than C++, and so there is no chance that C++ takes over as the lingua Franca of machine learning until those activities are rare
many ML experts are not software engineers. They just want syntax to get their job done. fair enough.
But memory wise, there is almost no optimization or reuse (and it's sacrificed for performance), which leads to insane memory usage.
And it's not that they are bad - but optimal compilation of a graph is combinatorially explosive problem, impossible without heuristics and guesswork (what to reuse and waste memory vs recompute). A good programmer can do a significantly better job.
Exploring the landscape ends up with you having 29384232938792834234 different python environments, because that one thing requires specific versions of one set of libraries, while that other thing requires different versions of the same library and there is no middle ground.
It's horribly annoying and I absolutely love python!
Python opened up the world of code to a lot more people, but there was a cost to that, and the real action as far as actual computer systems go is always gonna be at a much lower level, in much the same way that a lot of people can top up their fluids but most of us pay someone to change the oil.
I could actually totally see oil changes becoming completely robotic in the future, but we would first have to establish open standards for oil pans that all automakers adhered to.
The whole computers designing computers thing, outside of someone cracking cold fusion I don't think we'll ever have the juice for it. In my lifetime, a C-level programmer will never be out of work, but I suspect that the demand for Python programmers is going to slack off, while the supply continues to grow.
This project gained popularity due to the HIGH DEMAND for running large models with 1B+ parameters, like `llama`. Python dominates the interface and training ecosystem, but prior to llama.cpp, non-ML professionals showed little interest in a fast C++ interface library. While existing solutions like tensorflow-serving [1] in C++ were sufficiently fast with GPU support, llama.cpp took the initiative to optimize for CPU and trim unnecessary code, essentially code-golfing and sacrificing some algorithm correctness for improved performance, which isn't favored by "ML research".
NOTE: In my opinion, a true pioneer was DarkNet, which implemented the YOLO model series and significantly outperformed others [2]. Same trick basically like llama.cpp
[1] https://github.com/tensorflow/serving [2] https://github.com/pjreddie/darknet
Not to mention, to really protect source code and all the sugar around the training systems, it's going to be a good investment to get out of hobby land and manage your own memory and just code them in C/C++.
It strikes me that the hobby AI ethos aligns very well with scripting languages in that they both assume the availability of endless resources to push things a little dirtier and messier and see if anything interesting emerges. Which is great for hobby AI. It's probably not the future, though, unless resource availability outpaces the imagination of people to write more and more bloated scripts to accomplish what's already been proven.
Have you ever tried to iterate developing any serious class of algorithms in C++?
> Once you know what you actually want it to do
When is that exactly? Even the last few months of LLM land development show very clearly how everything is rapidly evolving (and will very likely continue for quite some time).
Numerical linear algebra stabilized decades ago so you do have low-level libraries in C++ (or even fortran) but there is quite some distance between an LLM and linear algebra.
Yes, and it's not so bad. A lot of ML deployments have been based on C/C++ for inference anyway (with Python driving the training). So that's really nothing new. I.e. most Python research code is not deployable in terms of quality / performance.
In the absence of alternatives its workable. Libraries like eigen or armadillo help a lot. But Python with numpy has been extremely popular for a good reason.
There wasn’t a python library like that anyone was used to using. Would have always been a C extension anyway.
Starting with the cpu in this situation made sense, strangely. There are python wrappers now. I tried to make one for ya’ll in rust in April, but haha I had a compiler issue I never solved.
Given the author hand-rolled his own FFT, I'm also guessing it's not as performant?
In comparison I could just type git clone https://github.com/ggerganov/llama.cpp and make . And it worked. And since then I've managed to get llama.cpp clBLAS partial GPU acceleration working with my AMD RX 580 8GB. Plus with the llama.cpp CPU mmap stuff I can run multiple LLM IRC bot processes using the same model all sharing the RAM representation for free. Are there even ways to run 2 or 3 bit models in pytorch implementations like llama.cpp can do? It's pretty rad I could run a 65B llama in 27 GB of RAM on my 32GB RAM system (and still get better perplexity than 30B 8 bit).
You're comparing a single, well managed project that had put effort into user onboarding against all projects of a different language and proclaiming that an entire language/ecosystem is crap.
The only real take away is that many projects, independent of language, put way too little effort towards onboarding users.
Python is a mess. llama.cpp was literally a git clone followed by "cd llama.cpp && make && ./main" - I can recite the commands from memory and I haven't done any C/C++ development in a long time.
The example you gave works because llama.cpp specifically strives to have no dependencies. But this is not an intrinsically useful goal; there's a reason software libraries were invented. I always have fun when I find out that the thing I'm trying to compile needs -std=C++26 and glibc 3.0, and I'm running on a HPC cluster where I can't use a system level package manager, and I don't want to be arsed to dockerize every small thing I want to run.
For scientific and ML uses, conda has basically solved the whole "python packaging is a mess" that people seem to still complain about, at least on the end-user side. Sure, conda is slow as hell but there's a drop in replacement (mamba) that solves that issue.
Coming from rust (and nodejs before that), the package management situation in python feels like a mess. It’s barely better than C and C++ - both of which are also a disaster. (Make? Autotools? CMake? Use vendored dependencies? System dependencies? Which openblas package should I install from apt? Are any of them recent enough? Kill me.)
Node: npm install. npm start.
Rust: cargo run. Cargo run —-release.
I don’t want to pick from 18 flavours of “virtual environments” that I have to remember how to to “activate”. And I don’t want to deal with transitive dependency conflicts, and I don’t want to be wading through my distro’s packages to figure out how to manually install dependencies.
I just want to run the program. Python and C both make that much more difficult than it needs to be.
In general solving such environments (where some versions are left floating or only have like >= requirements) is an NP-hard problem. And it requires care and has to draw source code and/or binaries from various sources.
This is the problem that conda/mamba solves.
If you just want to install python packages, use pip+virtualenv. It's officially supported by python. And while pip has traditionally been a mess, there's been a bunch of active development lately, the major version number has gone from like 9.xx to 23.xx in like the past two or three years. They are trying, better late than never, especially for an ecosystem as developed as python's.
So, if you want to compare rust/cargo, and it handles non-rust deps, then the equivalent is conda. Otherwise, it's virtualenv+pip. I don't think there are any other serious options, but I agree that two is not necessarily better than one in this case. Not defending this state of affairs, just pointing out which would be the relevant comparison to rust.
You can run arbitrary code at build time, and you can link to external libraries. For the likes of C libraries, you’ll typically bind to the library and either require that it be installed locally already, or bundle its source and compile it yourself (most commonly with the `cc` crate), or support both techniques (probably via a feature flag). The libsqlite3-sys crate is a fairly complex example of offering both, if you happen to want to look at an actual build.rs and Cargo.toml and the associated files.
> [pip’s] major version number has gone from like 9.xx to 23.xx in like the past two or three years.
It skipped from 10.0.1 to 18.0 in mid-2018, and the “major” part has since corresponded to the year it was released in, minus 2000. (Look through the table of contents in https://pip.pypa.io/en/stable/news/ to see all this easily.)
Yeah, no… my current $DAYJOB has a confusing mix of nx, pnpm, npm commands in multiple projects. Python is bad but node is absolutely not a good example.
I understand why people find Python's packaging story painful, I guess is what I'm saying.
mamba can do this much quicker.
No need to denigrate those who don't need to import all their code; it's perfectly fine to not rely on third parties when developing software - some of our best works result from this.
I know exactly what you mean, but I'm probably so inured to it by now that I've just come to accept it. Obviously not everyone feels this way!
It's annoying to see them "glueys" compared to the binary compiled segment where the heavy lifting is done. Python and others exist to latch on and assimilate. Resistance is futile:
https://pypi.org/project/pyllamacpp/
https://www.npmjs.com/package/llama-node
A self contained zero dependency Python project is also literally “python app.py”
The barriers to entry are lower with python, and so I think there tends to be a lower standard for deployment.
Deploying a PyTorch running Python app to anything in a way a user can just run it is a struggle. Even iOS aside that’s not a small use case that’s all local and offline ML potential.
Really makes me wonder about the whole language, did they ever expect the code to have to run elsewhere than the machines the writer controls.
The desktop/laptop LLMs use case is one where resource efficiency often makes the difference between “I can’t use this at all” rather than the more frequent “I can use it but it maybe runs a little slower”, and llama.cpp offers that. It also has offered new quantization options that the Python-based tooling hasn’t, which compounds the basic resource efficiency point.
(It’s also not an “everyone” thing: plenty of people are using the Python-based toolchains for LLMs, its possible for their to be multiple popular options in a space . Not everything is all or nothing.)
For LLMs, unless you're doing extensive work refactoring the inputs, there are fewer productivity gains to be had around the edges - the main gains are just speeding up training, evaluation and inference, i.e. pure performance.
Yes.
I have been playing around with whisper.cpp; it's nice because I can run the large model (quantized to 8-bits) at roughly real-time with cublas on a Ryzen 2700 with a 1050Ti. I couldn't even run the pytorch whisper medium on this card with X11 also running.
It blows me away that I can get real-time speech-to-text of this quality on a machine that is almost 5 years old.
I like Aiko for on-device transcription both in macOS and iOS https://apps.apple.com/us/app/aiko/id1672085276
https://twitter.com/ggerganov/status/1666120568993730561?s=4...
Godspeed
We're at the very dawn of this technology going mainstream, and you're saying that it's unlikely for new players to release new, competing and incompatible models?
OpenCL is (was?) the main open alternative to CUDA and was mainly backed by AMD and Apple. Apple got bored of it when they decided Metal was the future. AMD got bored of it when they developed rocm and HIP (basically a partially complete compatibility layer with CUDA).
There's also stuff like DirectML which only works on windows and e.g. various (Vulkan, directx etc.) compute shaders which are really more oriented at games.
There's also a bit of a performance aspect to it. Obviously GPGPU stuff is massively performance sensitive and CUDA gets you the best performance on the most widely supported platform.
AMD are also the main competitor in the space but all but totally dropped support for GPGPU in desktop cards, trying to instead focus on gaming for Radeon and compute in MI/CDNA. They seem to have realised their mistake a bit and are now introducing some support for RDNA2+ cards.
There was an implicit assumption that just building for more ISAs was no big deal. That assumption was wrong, but the good news is that big improvements to compatibility can be made even for existing hardware just by more thoughtful handling of the GFX ISAs.
If you know what you're doing, it's possible to run ROCm on nearly all AMD GPUs. As I've been packaging the ROCm libraries for Debian, I've been enabling support for more hardware. Most GFX9 and GFX10 AMD GPUs should be supported in packages on Debian Experimental in the upcoming days. That said, it will need extensive testing on a wide variety of hardware before it's ready for general use. And we still have lots more libraries to package before all the apps that people care about will run on Debian.
I don't understand why isn't SYCL more widely used.
Apple backed OpenCL because they needed an alternative after their divorce with Nvidia. No one was going to target an AMD alternative when they had such a trivial market share, so it had to be an open standard. Initially this arrangement was highly productive and OpenCL 1.x enjoyed terrific success. Vendors across compute markets piled support behind OpenCL and many even started actively participating in it. However this success is what precipitated the disastrous OpenCL 2.x series. In other words, OpenCL 2.x was far too revolutionary for many and far too conservative for others. What followed was Apple pulling out to pursue Metal, AMD having shoddy drivers, Nvidia all but ignoring it, and mobile chip vendors basically sticking to 1.2 and nothing more. Eventually this deadlock was fixed after OpenCL 3.0 walked back the changes of 2.x, but this was in large part because the backers of 2.x moved to SYCL.
As for AMD, OpenCL was a tremendous boon when it was first introduced. At least initially it gave them a fighting chance against CUDA. But it was never realistic for OpenCL to be a complete CUDA alternative. I mean, any standard that is basically “everything in CUDA and possibly more” is a standard no one could afford or bother to implement. ROCm and HIP are basically AMD using an API people are already familiar with software underneath to play to the strengths of their hardware.
“AMD are also the main competitor in the space but all but totally dropped support for GPGPU in desktop cards, trying to instead focus on gaming for Radeon and compute in MI/CDNA.”
Keep in mind that AMD has been under intense pressure to deliver world class HPC systems, and they managed to do so with ORNL Frontier. I don’t blame them for being selective with ROCm support because most of those product lines were in flight before ROCm started development in earnest. That said, Nvidia is obviously the clear leader for hardware support, as therefore the safest option for desktop users.
So… partial compatibility layer?
In this case I still prefer my version. I feel that it puts greater emphasis on the fact that it can potentially be complete, given the massive value that could give to the project.
Also "I would have written you a shorter letter but I did not have the time" sentiment springs to mind.
(I note that if we had an "open" GPU architecture in the same way that we have CPU architectures, things might be a lot better, but the openness of the IBM PC seems to be a historical accident that no company will allow again)
Now that AMD has the capital they are playing catch up to Nvidia. But it’s going to take time for their software to improve. Hiring at boat load of programmers all at once isn’t going to solve that.
That being said, it's important to note that ROCm is Linux only. Not only that, but ROCm's GPU support has actually been decreasing over the past few years. The current list: https://rocm.docs.amd.com/en/latest/release/gpu_os_support.h... Previously (2022): https://docs.amd.com/bundle/Hardware_and_Software_Reference_...
The ELI5 is that a few years back, AMD split their graphics (RDNA) and compute (CDNA) architectures, which Nvidia does too, but notably (what Nvidia definitely doesn't do, and a key to their success IMO) AMD also decided they would simply not support any CUDA-parity compute features on Windows or their non "compute" cards. In practice, this means that community/open-source developers will never have, tinker, port, or develop on AMD hardware, while on Nvidia you can start with a GTX/RTX card on your laptop, and use the same code up to an H100 or DGX.
llama.cpp is a super-high profile project, has almost 200 contributiors now, but AFAIK, no contributors from AMD. If AMD doesn't have the manpower, IMO they should simply be sending nsa free hardware to top open source project/library developers (and on the software side, their #1 priority should be making sure every single current GPU they sell is at least "enabled" if not "supported" in ROCm, on Linux and Windows).
[1] https://github.com/ggerganov/llama.cpp/issues/1433#issuecomm...
https://geohot.github.io/blog/jekyll/update/2023/06/07/a-div...
I think they missed the thrust of the rant.
Firstly, nVidia has been at it much longer. Just because of this tools on nVidia side feel easier to set up / are more polished (at least that was my feeling when fiddling with ROCm like a year ago).
Second but still related to #1, from the beginning even consumers nVidia cards were able to run CUDA and this made so that hobbyist and prosumers/researchers on a budget bought nVidia cards compounding even further the time/tooling advantage nVidia had. I.e. a huge user base of not only gamers but people that use their cards to do other things than gaming and know that things work on these cards.
These are, IMHO, the main reasons why everyone targets CUDA and explain why frameworks like Tensorflow or Pytorch targeted it as a first class citizen.
But the reality is that when Tensorflow and Pytorch came to be there was no alternative. Now you need to jump through hoops to make it work with non CUDA hardware.
Additionally, while drivers play a role, I think the main difference is in the computing libraries (CUDA vs ROCm)
Someone who knows about this space that can comment on the likelihood that SYCL will be a good option eventually? Cross platform and cross vendor compatibility would be really nice, and not supporting the proprietary de facto standard would also be a bonus as long as the alternative works well enough.
[0] https://www.khronos.org/sycl/
[1] https://spec.oneapi.io/versions/latest/elements/sycl/source/...
I think that's the problem. Khronos isn't known for good UX, and being from Khronos is exactly the reason why I'm not even bothering to check it out. I want an alternative to CUDA, but I also want it to be as easy to use as CUDA.
(Vulkan is from 2015 and is just recently starting to become usable.)
They're supposed to be more optimized and more stable compared to AMD. That's how it was before anyway, not sure today.
On top of that, AMD's compute stack is fairly immature, their OpenCL support is buggy and ROCm compiles device specific code, so it has very limited hardware support and is kind of unrealistic to distribute compiled binaries for. Then, getting to the optimization aspect, NVIDIA has many tools which provide detailed information on the GPU's behavior, making it much easier to identify bottlenecks and optimize. AMD is still working on these.
Finally, NVIDIA went out of its way to support ML applications. They provide a lot of their own tooling to make using them easier. AMD seems to have struggled on the "easier" part.
Maybe someone more clever than me can go into the specifics, I only understand the minimum of the low lvl GPU details.
Nice high lvl document
[0] https://www.acecloudhosting.com/blog/cuda-cores-vs-tensor-co...
So, while I can't talk about the hardware differences in detail, developer experience is greatly on nVidia side and now AMD has a moat to overcome to catch up.
[edit] realized it's relative to the merged partial CUDA acceleration, so the speedup is more impressive, but still surprised by the core usage.
because the core's job is solely to direct the GPU, which is doing all of the work.
They are also innovating (or at least implementing innovations from papers) different ways to fit bigger models in consumer HW, making them run faster and with better outputs.
Pytorch and other libs (bitsandbytes) can be horrible to setup with correct versions, and updating the repo is painful. PyTorch projects require a hefty GPU or enormous CPU+RAM resources, while llama.cpp is flexible enough to use GPU but doesn't require it and runs smaller models well on any laptop.
ONNX is a generalized ML platform for researchers to create new models with ease. Once your model is proven to work, there are many optimizations left on the table. At least for distributing an application that relies on LLM it would be easier to add llama.cpp than ONNX.
Also, llama.cpp is excellent at splitting the workload between CPU and accelerators since its so CPU focused. You can run 13B or 33B in a 6GB GPU and still get some acceleration.
Also, as said above, quantization. That is make or break. There is no reason to run a 7B model at fp16 when you can run 13B or 30B in the same memory pool at 2-5 bits.
So basically GGML will run on hardware with less memory.
I’m curious how it compares to his apple silicon numbers
Heck, even using CPU, I've been able to generate 512x512.
If you arent generating 512x512 pictures, lmk, I'll go grab my bat file's startup parameters.
Just to be clear, I can also make 512x512 with CPU, so you basically just need to have the correct config of the .bat file.
https://huggingface.co/TheBloke/open-llama-7b-open-instruct-...
It's stupidly simple to get going.
If this is faster than torchlib this optimizations should flow to torchlib as well
Love the idea of not having to deal with python, though; dependency management is just horrible, I'd much rather have ML projects written in cpp.
I find the cost of computing extremely affordable, even for high end stuff. Whats the amortization on a 2-3k computer over 7 years? How about if I use it 4 hours a day actively and 24 hours passively?
I have considered spending 10-30k on a computer given the recent AI craze, but the thing stopping me is that by 2025, a 10-30k computer in the AI space is going to be 2-4x better. Only in the last 1 year are we finding out the importance of absurd amounts of VRAM. I feel like the 4090's 24gb VRAM is going to age alright at best, but most likely poorly. (Not that 4090 buyers are going to have qualms upgrading to the 6090)
I suppose that's it for now.
Oh, I thought of another use, I run a small business on the side and my interns occasionally don't have a laptop, I give them a crappy laptop. (they are basically just using excel/google sheets)
it runs on the CPU
this news story is that they are now extending GPU support to llama.cpp
You can probably find a google colab link.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
This user has some models already compiled for use with GGML (look for models with that in the name):
https://huggingface.co/TheBloke
Or if you want to convert your own model the llama.cpp repo has good instructions. Briefly it's `python3 convert.py <model>` - then if you are using a large parameter model you may need to quantize it to fit in memory `./quantize <source_model> <destination_name> <quantization>`
llama.cpp specifically has Falcon on their roadmap, and some other quantized implementations already work with it. But the transition will be slow.
amen.
A lot of pissant shitty defenses of monopolization too. Wrong or right, this is a shit world we're in now. https://news.ycombinator.com/item?id=36304225
As LLM adoption grows, I wonder whether Apple's approach will start to make more sense for consumer adoption, so that you can run models on your machine without needing to pay large subscription costs for AI-powered apps (since OpenAI et al have fairly high fees). The high cost of using the APIs in my opinion is a drag on certain types of adoption in consumer apps.
Llama.cpp actually first started as a way to run LLMs on Macs! At first CPU-only, but then later the first GPU driver backend added was Metal, not anything from Nvidia.
The PC gamers really want the GPU component to be separately upgradable from the CPU. Non-gamer PC users don't care about the GPU performance, just cost. The datacenter folks want to be able to use a single $1k CPU to host $100k worth of GPUs.
It's plausible that AI accelerators follow a different path for the consumers. It's harder to see it happening for the datacenter market.
Being able to switch from an optimized CPU-centric workload to an optimized GPU-centric workload without any hardware changes sounds useful to me.
You could even do unified memory with upgradeable separate CPU and GPU. You won’t get the benefits of having them on the same chip, but there’s nothing intrinsic about the separation requiring separate memory space.
NVIDIA never sold a GPU with expandable memory, either.
A100 goes up to 2TB/s in the largest configuration, and H100 claims 3TB/s per GPU node. (These figures keep changing with new variants.) But you can buy several Mac Studios for the price :-)
The real use of H100 is for training, as it can pool many GPUs together with a dedicated high bandwidth network. You can't do that with Mac Studios.
Since the models are not suppressed to answer certain questions and indeed have been demonstrated to be biased toward one end of the political spectrum, and if you treat the model’s output as a knowledge graph, as proposed by John Schulman, one of the cofounders of OpenAI, then yes, I would say the suppression of freedom of speech is a valid argument to make. Otherwise why would there be a set of “uncensored” models that exist in the open source world?
I would suggest you read about these models and think about the implications. Perhaps that’ll lead to you to reconsider your stance
OpenCL was the utopia timeline.
It seems like a no brainer to hire a bunch of people to work on making PyTorch / Tensorflow on AMD become a competitive option. It’ll just take a few years.
What GP seems to be asking for is an open CUDA replacement, which is kinda like asking someone to fund a Free and Open Source cruise ship to compete with Carnival for you. You'll get somewhere with some effort, luck and good old human intuition, but Nvidia can outspend you 10:1 unless you have funding leverage from FAANG.
It turns out you need very different kernels for good performance on different GPUs, so OpenCL is a nice tool, but not sufficient; you need a hardware-specific kernel library.
From the framework side, each integration is relatively expensive to support, so you really don’t want to invest in many of them. Without some sort of kernel API standard, you’re into a proprietary solution, and NVidia did an amazing job at investing in their software, so that’s the way things go.
I think we had a pretty solid foundation for doing something smarter with PlaidML, but after we were bought by Intel, some architectural decisions and some business decisions consigned that to be a research project; I don’t know that it’s going anywhere.
These days, I’d probably look into OctoML / TVM, or maybe Modular, for a better solution in this space… or just buy NVidia.
(I worked a bit on Intel’s Meteor Lake VPU; it’s a lovely machine, but I’m not sure what the story will be for general framework integrations. I bet OpenVINO will run really well on it, though :-)
Well, at least you admit it's not Nvidia's fault. Apparently Apple, Intel and AMD don't think there's much money to grab here.
It stands to reason that Intel is making highly price-competitive hardware at the moment, but people don't talk about them as much Nvidia because they have a minuscule install base with primitive Windows drivers. I wouldn't count them out if their first showing is this impressive, though.
But AMD is coming strong, and they are trying to compete with Nvidia now. https://www.forbes.com/sites/iainmartin/2023/05/31/lisa-su-s...
Think of Apples ecosystem vs Android or MS's Office-Outlook-Teams vs anything else.
Disclaimer: I work at Qualcomm.
This processor shows up in both snapdragon SoCs and cloud ai 100.
But I think you're getting downvoted since it's very off-topic.
I suspect it would be more of a legal challenge than a technical one.
exllama is really memory efficient and really fast
[0] https://huggingface.co/docs/transformers/main/model_doc/llam...
[1] https://github.com/turboderp/exllama
[2] https://github.com/Lightning-AI/lit-llama
EDIT: Or do you mean cuda? Because yeah, it's such a shame AMD's Rocm is so bad even geohot gave up. it's examples don't even run without crashing.
https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuec...
edit: Note that this is my project.
AMD gave him a binary blob driver and that fixed his problem. Also, tinygrad is the only Python framework I know that has full OpenCL acceleration.
So far it doesn’t look that AMD is fully on board with Tiny Corp, but they are talking…
a) CUDA won in a free market because NVidia showed they cared about it
b) Llama has support for OpenCL (via CLBlast) and Apple Metal
The OpenCL support already has a custom kernel for token generation.
No open source though.