Blaze: A High Performance C++ Math Library
bitbucket.org
bitbucket.org
EDIT: also being on Sourceforge is kind of a hinderance to discovery these days. I wonder why they chose to be on there instead of github?
I have had very little success with most of the open self-hosted ones, even with my 4xA40 setup, as they either don't know the c++ libraries or generate very good-looking numpy stuff, full of horrors, simple and very very subtle bugs...
Looking for the same thing from any linear algebra library or language to cuda BTW (yes, calls to cu-blas/solver/sparse/tlass/dnn are OK), I haven't found one model able to write cuda code properly - not even kernels themselves but at least chaining library calls.
Probably doesn't exist (invoking Cunningham's Law).
Large amounts of repetitive yet meaningfully detailed code. Algorithms that can (and often are) implemented using different conventions or orders of operations. Edge cases out the wazoo.
A solid start seems like it would be using LLMs to write extensive test suites which you can use to verify these new implementations.
And yes, it's nice to build unit test and benchmark harnesses. But those were never really such time-wasters for me.
https://conradsanderson.id.au/pdfs/sanderson_curtin_armadill...
In general I would be skeptical about any benchmark that claims to beat MKL significantly on standard operations
For small matmul there is libxsmm. It may take tremendous efforts make something faster than oneDNN and libxsmm, as jit-based approach of https://github.com/oneapi-src/oneDNN/blob/main/src/gpu/jit/g... is too flexible: if someone finds a better sequence, oneDNN can reuse it without major change of design.
But MKL is not limited to matmul, I understand it...
Eigen uses C++ templates to do most things, which explodes compile times.
The reality is that almost all workloads aren't anywhere near saturating the AVX instruction max bandwidth on a CPU since Haswell.
That’s true, but GPUs aren’t only good at FLOPs, the memory bandwidth in them is also an order of magnitude faster than system memory.
In my previous computer, the numbers were 484 GB/second for 1080 Ti, and 50 GB/second for DDR4 system memory. In my current one, they are 672 GB/second for 4070 Ti super, and 74 GB/second for DDR5 system memory.
I imagine we'll get to a point where CPUs are actually just pretty dumb drivers for issuing gpu commands.
Did GPUs win?
It's a fairly fractal pattern in distributing computing. Move the high throughput heavy computation bits away from the low latency responsive bits ("low latency" here is relative to the total computation). Use an event loop for the reactive bits. Eventually someone will invert the event loop to use coroutines so everything looks synchronous (Go, anyone? python's gevent?).
After it seems to me that the only real question is if takes too long or costs too much to move the data to the storage location the heavy computation hardware uses. There's really not much of a conceptual difference between airflow driving snowflake and c++ running on a cpu driving cuda kernels. It takes a certain scale to make going from a OLTP database to an OLAP database worth it, just like it takes a certain scale to make a GPU worth it over simd instructions on the local processor.
There is great power in the convenience of "with open('foo') as f:". Most workloads are still stitching together I/O bound APIs, not doing memory-bound or CPU-bound compute.
It took a long time to find something that really took advantage of it, but we did eventually. CUDA enabled deep learning which enabled LLMs . That's history.
What surprised me about the statement was that it implied that the model of python driving optimized GPU kernels was broader than deep learning.
That was the original vision of CUDA - most of the computational work being done by massively parallel cores
GPUs are still very limited, even compared to the SIMD instruction set. You couldn't make a CUDAjson the same way the SIMDjson library is built for example, because it doesnt handle SIMD branching in a way that accomodates it.
Second, again, the latency issue. GPUs are only good if you have a pipeline of data to constantly feed it, so that the PCIe transfer latency issue is minimal.
The penalty for branching has reduced in the last years, but yeah it's still heavy, but if you're OK with a bit of wasted compute, you can do some 'speculative' execution and do both branches in different warps, use only one result...
But yes, you're still using an accelerator.
That a math library of all things could be complete is several orders of thinking beyond their ability. I'm sure the gut reaction is to downvote this for the embarrassing criticism, but in all seriousness, this is the right answer.
But you can't just call things "dead" for no reason, it's in poor taste. It's feature-complete, not dead!
What you're actually saying is you expect open source maintainers to add arbitrary functionality for free.
Sure, but the discussion here is about a software library not the math concepts
Short of a bug in the implementation, there has yet to be a valid explanation for why mathematics libraries need to be continuously maintained. If I published an NPM library called left-add, which adds the left parameter to the right parameter (read: addition) how long, exactly, should I expect to maintain this for others?
The only explanation so far is that scumbags expect open source library maintainers to slave away indefinitely. The further we steer into the weeds of ignorant explanations, the more I'm inclined to believe this really is the underlying rationale.
1: https://en.wikipedia.org/wiki/Curry%E2%80%93Howard_correspon...
1. Bug fixes
2. Security issues
3. Optimization
4. Compatibility/Adapt to landscape changes
People pointing flaws in a library aren't "scumbags that expect open source library maintainers to slave away indefinitely"
No one is forcing the maintainer to "slave away", they can step down any time and say I'm not up for this role anymore. Those interested will fork the library and carry the torch.
No need to be so defensive and insult others just for giving feedback.
Regardless of the strawman, the person(s) that authored the code don’t owe you anything. They don’t have to step down, make an announcement, or merge your changes just because you can’t read or comprehend the license text that says very clearly in all capital letters the software is warrantied for no purpose once so ever, implied or otherwise.
If one had a patch and was eager to see it upstreamed quickly, it seems like you’re arguing the maintenance status actually doesn’t matter. Since "[t]hose interested will fork the library and carry the torch" if the patch isn’t merged expediently.
But if you're confident the interested will fork and carry the torch, why do you think you're entitled to force the author(s) giving software warrantied for no purpose should step down. That's genuinely deranged, and my insults appear to be accurate descriptions rather than ad hominem attacks since no coherent explanation has been provided as to why the four reasons given somehow supersede the authors chosen license.
If you have a math library that is relying on hardware and compilers to make it fast you should acknowledge that the software and hardware ecosystem in which you exist is constantly changing even if the math is not.
This is a pretty bold and loud acknowledgement.
What more could you really ask for when even lawyers think this is sufficient.
Some signal that the project is being maintained? If it’s not that’s fine but don’t go radio silent and get pissy when people ask if a project is dead…
This is not a legal or moral issue it’s just being considerate for others as well. You, the maintainer, made the choice to maintain this project in the public and foster a userbase. This is not a one-way relationship. People spend their time making patches and integrating your software. You are under no obligation to maintain it of course but dont be a dick.
The implication is the mistake, not the author for not being explicit enough.
Further, what makes you assume everyone is on the same page about what that social contract is? Have you even considered the possibility that there might be differences of opinion on a social contract which are incompatible? It's why the best course of action is to follow the license rather than delusional fantasies.
The idea there's a social contract is sophistry. Plain and simple.
And I wish Eigen had a larger spectrum of 'solvers' you can chose from, depending on what you want. But in general I agree with you, except there's always a cycle to eke out somewhere, right?
The codebases I've seen in many game physics engines seem to all roll their own minimal math libraries for these stuff, or even just use SIMD (SSE / AVX) intrinsics directly. Examples: PhysX (https://github.com/NVIDIA-Omniverse/PhysX), Box2D (https://github.com/erincatto/box2d), Bullet (https://github.com/bulletphysics/bullet3)...
What is of import here?
Previous post (by the same submitter): https://news.ycombinator.com/item?id=34407106