Second, OP, like llama.cpp, produced efficient and highly specialized code after it was clear the model being specialized for (StableDiffusion / LLaMa / …) works well. Where Python shines, though, is the prototyping phase when you have yet to find an appropriate model. We have yet to see this sort of easy & convenient prototyping in C++.
Now, this is not to take away anything from the fantastic work that's being done by the llama.cpp people (to whom I also count OP) in the "ML on a CPU" space. But the problems being solved are entirely different.
+1.
To produce a highly-optimized C/C++ kernel that utilizes the CPU to the fullest extent, it requires tremendously amount of talent and expertise. For example, not everyone can write a hand-vectorized kernel with AVX2 intrinsics (outside a few specialized applications like 3D graphics, media encoding, and the likes), and even fewer people can exploit the underlying feature of the algorithm for optimization, such as producing usable output at greatly reduced numerical precision. The power of LLM provides strong motivation to drive the brainpower of countless programmers all over the world to do just that. New techniques are proposed and implemented on a monthly basis, with people thinking and applying every possible trick on the LLM optimization problems. In this regard, moving from Python to C is totally reasonable.
In comparison, right now I'm working on optimizing a niche open-source scientific simulation kernel with a naive C codebase. Before me, there were hardly any contributors in the last decade.
Python has its place because not everyone has a level of resource and expertise comparable to ML. In particular, when the bulk of the data processing of a Python script is in done in a function call to a C++ or FORTRAN kernel like scipy, the differences between naive C and naive Python code (or Julia code if you're following the trend) are not that much, especially when it's a one-off project for just publishing a single paper.
If your guys arent on this I'd suggest you get them on it, it dramatically simplifies setup
$ dvc pull
Command 'dvc' not found, but can be installed with:
sudo snap install dvc
$ sudo snap install dvc
error: This revision of snap "dvc" was published using classic confinement and thus may perform
arbitrary system changes outside of the security sandbox that snaps are usually confined to,
which may put your system at risk.
If you understand and want to proceed repeat the command including --classic.
ok I get dvc installed somehow -- don't remember. Time to get the weights...
$ python3 -m dvc pull
ERROR: unexpected error - Forbidden: An error occurred (403) when calling the HeadObject operation: Forbidden
Having any troubles? Hit us up at https://dvc.org/support, we are always happy to help!
Finally I just have my colleague manually copy the weights. This kind of thing went for hours.Thanks for giving DVC a try!
There are a few ways to install dvc, see https://dvc.org/doc/install/linux
With snap, you need to use `--classic` flag, as noted in https://dvc.org/doc/install/linux#install-with-snap Unfortunately that's just how snap works for us there :(
Regarding the pull error, it simply looks like you don't have some credentials set up. See https://dvc.org/doc/user-guide/data-management/remote-storag... Still, the error could be better, so that's on us.
Feel free to ping us in discord (see invite link in https://dvc.org/support). I'm @ruslan there. We'll be happy to help.
What even is dvc
edit: also- i'd avoid snap and just use your regular package manager.
In most cases, it would be possible to do close to that, but it is extremely common to run into things being distributed in the AI/ML space with install instructions that don’t include that, and instruct you to have a global install of a certain Python version, and then to pip install the dependencies (and globally install non-Python package dependencies, if there are any), so even if they’d work in a venv, you have to (1) indepently know you should be doing that, and (b) translate the instructions – which where (1) applies is usually trivial if all the dependencies are proper python packages, but can be more involved otherwise.
So, yeah, I can see that a lot of the time the path of least resistance is just to create an isolated container environment for it.
I feel like the person you're replying to knows that the GPU is better suited than the CPU to do this task, and your argument doesn't really make sense. I think they were referring to the python venv environment with all the library dependencies as the "specialized environment"
Obviously doing any heavy lifting in Python is a bad idea. But as a scripting language I think it's good, especially if you keep the environment simple. I don't think the answer for DL training is to dump Python entirely and start over in pure C/C++/Rust/Julia/whatever. Learning C/C++ is too big of an ask for everyone working on the model design and training side and it would slow down progress significantly - most of that work is actually data munging and targeted model tweaks. But I do think there's still a lot that can be done to decouple Python from the underlying engine and yield networks where inference can be run in a minimal dependency environment. There's lots of great people working on all these things.
CERN already used prototyping in C++, with ROOT and CINT, 20 years ago.
Nowadays it is even usable from Netbooks via Xeus.
It is more a matter of lack of exposure to C++ interpreters than anything else.
When was the last time you looked at llama.cpp? It has supported GPU, GPU+CPU, and distributed inference using OpenMPI for awhile now. It also supports training, as well as negative prompting and grammars! The ease of getting llama.cpp running on just about anything has already started innovation.
The problem with PyTorch specifically is that (without Triton compilation) pretty much all projects run in eager mode. That's fine for experimentation and demonstrations in papers, but its crazy that its used so much for production without any compilation. It would be like using debug C binaries for production, and they only work with any kind of sane performance on a single CPU maker.
In addition to compute the GPU architecture is one that somewhat colocates working memory alongside compute. Units have local memories that sync with global memory. Is that a big part of why GPUs are so good for this?
Yeah, sort of.
LLMs like llama at a batch size of 1 are hilariously bandwidth bound.
Stable Diffusion less so. Its still bandwidth heavy on GPUs, but compute is much more of a bottleneck.
Especially since the python training systems are mostly calls into libraries written in C++!
both Cuda and the Metal shader language are C++, so is OpenCL since 2.0 (https://www.khronos.org/opencl/), so is AMD ROCm's HIP (https://github.com/ROCm-Developer-Tools/HIP), so is SYCL (https://www.khronos.org/sycl/)? C++ is pretty much the language that runs most on GPUs.
> no vector instructions,
There's a thousand different possibilities for SIMD in C++, from #pragma omp simd, to libs such as std::experimental::simd (https://en.cppreference.com/w/cpp/experimental/simd/simd), Eve (https://github.com/jfalcou/eve), Highway (https://github.com/google/highway), Vc (https://github.com/VcDevel/Vc)...
C++ is one of the supported CUDA languages, even standard C++17 does run just fine on the GPU.
Metal uses C++14 alongside some extensions.
It’s technically not a C++ feature, but both gcc (https://gcc.gnu.org/onlinedocs/gcc/Vector-Extensions.html) and Clang (https://releases.llvm.org/3.1/tools/clang/docs/LanguageExten...) have vector types, and clang even supports the gcc way of writing them, so it gets pretty close.
Though it's certainly gotten better, the reason people push those is that they're written by compiler authors, who don't want to hear that their compiler doesn't work.
Some of the reason for this is that C doesn't let you specify memory aliasing as precisely as you want to. Fortran is better about this.
It really doesn't have any redeeming characteristics vs. Common Lisp, or Haskell, to warrant this bizarre popularity imo
I agree that its popularity is very odd, but academics take what they are given when attending fully paid conferences (aka vacations).
I think it would be very confusing for a child to start with a language so far away from low-level logic.
...And some people said BASIC was evil. At least what it is doing looks plain and direct.
Why?
I started with C++ and when they showed me C# I instantly feel in love cuz I didn't have to deal with unnecessary complexity and annoyances and could focus on pure programming, algorithms, etc.
but you are confirming my point :) ...You started with C++, then went to C#...
Both: high-to-low and low-to-high have some advantages, but it's not like one is always better than the other.
high-to-low allows you to write stuff earlier - like programs that do something useful, GUI, web, whatever.
but at the cost of understanding internals / under the hood.
On the other hand, it made me think a lot about computer memory, and it makes computer memory easy to work with. Now I’m really comfortable with memory and encodings so that’s nice. I don’t think I would have gotten that by starting with Java or Python.
Depending on the person.
For some, it would be very frustrating to start with a language so close to the implementation detail, and so far away from what you want to do. It's very possible that someone might have long lost the motivation before one can do anything non-trivial.
I started from Python, to C, to assembly, to 4-layer circuit boards. Whenever I went a level deeper, it feels like opening the inner working of a blackbox that I normally only interacts with pushbuttons on its front panel, but I otherwise is roughly aware of what they do.
On the other hand, much of my childhood was spent on tinkering with PCs and servers, including hosting websites and compiling packages from source, so I was already well aware of the basic concepts in computing before I started programming. So, top-down and bottom-up are both absolutely workable, under the right circumstances.
"Code Persona Attack"
Python is fine.
And they're right. Python is not a well designed programming language - it has exceptions and doesn't have value types so that's two strikes against it.
Of course, C++ isn't either.
If you don't know what you are doing, if you are exploring ideas, Rust will just get in the way. At some point you will end up realizing you need to adjust lifetimes, and that will require you to touch non-trivial amount of your code base. If you need to that multiple times, friction will overwhelm your desire to code.
I have a pet theory that, the people that find Rust intuitive and fun, are the people that are working on well beaten paths; Rust is almost boring at doing that, which is a good thing. And the people that find Rust gets in their way are the people that like to experiment with their solutions, because there aren't any set, trusted solutions within their problem space, and even if there are, they like to approach the problem on their own, for better or worse.
In any case:
> why would anyone start a greenfield project like this in C++ these days?
The video game industry can single-handedly carry C++ on their back, kicking and screaming, if need be. Rust is uniquely unfit to write gameplay code due to game development's iterative nature. Using scripting languages doesn't cut it either, because often, slower designer made scripts will need to be converted to C++ by a programmer, and pull in the crazy reference hell of the game state into the C++ land.
I would say Rust is OK for engine level features -- those don't change that often, and requirements are usually well understood. But that introduces a cadence mismatch between different systems too, so there is a cost there as well. But for gameplay? There's a reason why many Rust based game engines use crazy amount of unsafe Rust to make their ECS. Just not a good fit.
And of course, there's the consoles, where Sony seem to have a political reason for not supporting Rust on non-1st-party studios. I have no idea what they are thinking, honestly.
Does CUDA even have Rust bindings, and if so, are they on the same level as the C++ ones?
What do you mean by "the windows projects" that shift towards Rust?
Edit: I can't argue about the drama part. The competing compilers will get there. A couple gcc frontends in work, and crane lift as a competing back end for llvm and full self-hosting. There is also miri I guess to emit c? People use that to get rust on the C64 or other niche processors.
Meanwhile, Visual Studio team released better tooling for Unreal in Visual C++.
TLDR: quite often, using C++ instead of Rust saves software development costs.
Some software needs to consume many external APIs. Examples on Windows: Direct3D, Direct2D, DirectWrite, MediaFoundation. Examples on Linux: V4L2, ALSA, DRM/KMS, GLES. These things are huge in terms of API surface. Choose Rust, and you gonna need to write and support non-trivial amount of boilerplate code for the interop. Choose C++ (on Linux, C is good too) and that code is gone, you only need well-documented and well supported APIs supplied by the OS vendors.
Similarly, some software needs to integrate with other systems or libraries written in C or C++. An example often relevant to HPC applications is Eigen. Another related thing, game console SDKs, and game engines, don’t support Rust.
For the project being discussed here, GGML, for optimal performance the implementation needs vector intrinsics. Technically Rust has the support, but in practice Intel and ARM are only supporting them for C and C++. Not just CPU vendors, when using C or C++ there’re useful relevant resources: articles, blogs, and stackoverflow. These things help a lot in practice. I don’t program Rust, but I program C# in addition to C++, technically most vector intrinsics are available in the current version of C#, but they are much harder to use from C# for this reason.
All current C and C++ compilers support OpenMP for parallelism. While not a silver bullet, and not available on all platforms supported by C or C++, some software benefits tremendously from that thing.
Finally, it’s easier to find good C++ developers, compared to good Rust developers.
On the last point, I will again assert that a good c++ developer is just a good rust developer minus a month of ramp, that you'll get nack from not having to fight combinations of automake, cmake, conan, vcpkg, meson, bazel and hunter.
About SIMD, automatic vectorizers are very limited. I was talking about manually vectorized stuff with intrinsics.
I've been programming C++ for living for decades now. Tried to learn rust but failed. I have an impression the language is extremely hard to use.
Yes, rust directly supports modern intrinsics, that is what rustfft for instance uses. I try to stick with autovec myself, because my needs are simpler such that a couple tweaks usually gets me close to hand-rolled speedups on both avx 512 and aarch64. But for more complicated stuff yeah, rust seems to be keeping up. Some intrinsics are still only in nightly, but plenty of major projects use nightly for production, it is quite stable and with a good pipeline you'll be fine.
I've written c++ since ~94, and mostly c++17 since it came out. About a quarter of a century of that getting paid for it. I never liked or used exceptions or rtti, and generally used functional style except for preallocation of memory for performance. I think those habits might have made the transition a little easier, but the people on my team who had used a more OOP style and full c++ don't seem to have adapted much more slowly if at all. I struggled for years to internalize rust at home until I just jumped in at work by declaring the project I lead would be in rust. I have had absolutely no regrets. It really isn't as bad a learning curve as c++. But we learned c++ one revision at a time. Also, much like c++ rust has bits you mostly only need to know for writing libraries. So getting started you can put those things to the side for a bit right at first.
The crate you have linked contains just a single line of source code. Here it is:
pub struct SourceReader {}
Media foundation API reference: https://learn.microsoft.com/en-us/windows/win32/medfound/med...> rust directly supports modern intrinsics
C# also directly supports them, but it doesn’t help with usability. The support alone is not enough, the API needs to match the C compiler extensions defined decades ago by Intel and ARM.
Is it just a case of you forgetting how hard C++ was to learn?
I agree C++ is very hard to learn if you only have experience with higher-level languages like Python and Scala. I think there’re two reasons for that.
C++ is unsafe. There’s no way around this one, it was designed that way, like C or assembly. Still, with modern toolset it’s not terribly bad. Compilers print warnings, BTW I typically ask them to treat warnings as errors to deliberately fail the build. On Windows, a combination of debug build, debug C runtime, and visual studio debugger helps tremendously. Linux compilers have these sanitizers (address, memory, thread, undefined behavior) which are comparable, they too sacrifice runtime speed for diagnostics and debuggability.
Another reason, the language itself is very complicated, especially the templates. However, just because something is in the language doesn’t mean it’s a good idea to use it. You don’t need to be familiar with that stuff unless doing something very advanced, like customizing the Eigen C++ library. Don’t follow the patterns found in the standard library: unlike your code, that library has good reasons to use that template BS. If instead of templates you do something else, C++ becomes much easier to use, and most importantly other people will still be able to read and understand your code. Another reason to avoid excessive template metaprogramming, it slows down the compiler, because template-heavy code often needs to be in headers as opposed to cpp files.
P.S. If you don’t need extreme levels of performance (defined as “approach the numbers listed in CPU specs”, the numbers are FLOPS or memory bandwidth), and you don’t need the ecosystem too much, consider C# instead of C++. Much faster than Python, often faster than Scala or Java, easy integration with C should you need that (about the same as Rust, much easier than Python or Java), the only downside is these ~100MB of the runtime. The reputation is weird, but technically the language and runtime are pretty good. For example, here’s a C# library which re-implements a subset of ffmpeg and libavcodec C libraries for one particular platform, Linux on Raspberry Pi4: https://github.com/Const-me/Vrmac/tree/master/VrmacVideo
I suspect that if you really spent time learning rust that you'd really appreciate it coming from C++.
I use C++ for 2 main reasons, integration with other software or libraries, and performance.
When I need the performance, I don’t write conventional C++. I implement custom data structures, usually use vector intrinsics, sometimes use other platform intrinsics like BMI2, sometimes implement custom threading strategies on top of OpenMP or other platform-supplied thread pools. Most of these things are impossible or very hard to express in idiomatic safe Rust, which means the language constantly gets in the way.
Also, most programs contain both performance critical and performance agnostic pieces. When both pieces have non-trivial complexity, I typically use both C# and C++ for the software. Often the frontend part is in .NET, and the performance-critical backend is in C++ DLL.
Agility SDK and XDK have zero Rust support. If it isn't on the Agility SDK and XDK, it isn't official.
Hardly the same as the official Swift bindings to Metal, written in Objective-C and C++14.