https://developer.apple.com/documentation/accelerate/sparse_...
Accelerate is highly performant on Apple hardware (the current Intel arch). I expect Apple to ensure same for their M-series CPUs, potentially even taking advantage of the tensor and GPGPU capabilities available in the SoC.
By the way, if anyone at Apple reads this, thanks for the library, but, you know, calling conventions, algorithm, and options would really help on pages like this:
https://developer.apple.com/documentation/accelerate/sparsef...
Start here: https://developer.apple.com/documentation/accelerate/solving... and also watch the WWDC session from 2017 https://developer.apple.com/videos/play/wwdc2017/711/ (the section on sparse begins around 21:00).
There is also _extensive_ documentation in the Accelerate headers, maintained by the Accelerate team rather than a documentation team, which should always be considered ground truth. Start with Accelerate/vecLib/Sparse/Solve.h (for a normal Xcode install, that's in the file system here):
/Applications/Xcode.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX.sdk/System/Library/Frameworks/Accelerate.framework/Frameworks/vecLib.framework/Headers/Sparse/Solve.hMy sense is that Apple's focus is less on scientific computing and more so on enabling developers to build computation-heavy multimedia applications.
Your complaint is kind of strange. You're blaming "GPL restrictions" but the cost is for a commercial license.
And Rosetta will probably be around for a while...
It sounds like Intel produces an implementation of this thing that works on Intel and makes it available for free, whereas ARM don't (although another comment suggests Apple actually do), so you have to buy an expensive third-party implementation instead. That's not a difference that'll go away in the short term, and you can see why a processor company might legitimately choose one or the other approach.
> I'd have to explain to my commercial clients that there's another $50-100k upcharge due to the architecture change and licensing costs due to GPL restrictions.
If I'm a travel agent and an affordable hotel near a travel destination closes down, I might have to book my clients in a nicer but more expensive hotel. Their trip will be a bit more expensive. Or maybe they'll travel to a different city. It doesn't mean I dislike the nicer hotel.
I use GPL code all the time at home and I would license many things GPL, but there's no reason to push GPL software at corporations. They should have limited options and spend money, possibly expanding MIT code, possibly just raising the price of engineers by keeping engineers occupied.
> MKL has faster routines and is completely free, but it won't work on ARM
If your complaint is boo hoo, some people charge for software...well consider me unsympathetic.
That's really the point I'm trying to make and not to criticize anyone for using a GPL license. Moving to these new chips, in many cases, will be a much larger cost to an organization than just the cost of the computer.
Compared to not paying for it? Yes.
>Especially ironic because I'm sure the parent isn't working for free.
So? Who said that when you get paid yourself it stops being awful to have to pay for things?
- Here is a quote for 100k for adding SuiteSparse to the code.
- 100k‽ But I have found on the internet that SuiteSparse is free! Justify your quote.
At that point, they will have to explain to the client what GPL is and why they cannot use the free version.
I'm curious do people in numerical specialties say "codes" (instead of "code")? I don't often hear it that way but I'm not in that specialty.
I was trying to identify when, in normal usage, you'd say "numerical codes" rather than "numerical software" or just "numerical code". It seems a bit slippery!
Some contexts where it's prevalent: supercomputing, Fortran, national labs, large or multifaceted software. I also associate it with manager-speak ("our team has ported 77% of the simulation codes to HPSS").
software => codes
It can be compiled to use MKL, MUMPS, or SuiteSparse if available, but also has its own implementations. So you could easily use it as a wrapper to give you freedom to write code that you could compile on many targets with varying degree of library support.
Sadly, the factorization I personally need the most is a sparse QR factorization and PETSc doesn't really support that according to their documentation [1]. Or, really, if anyone knows a good rank-revealing factorization of A A'. I don't really need Q in the QR factorization, but I do need the rank-revealing feature.
[1] https://www.mcs.anl.gov/petsc/documentation/linearsolvertabl...
If you're a heavy user of SuiteSparse and upset about the license, you might want to check out Catamari (https://gitlab.com/hodge_star/catamari), which is MPLv2 and on-par to faster than CHOLMOD (especially in multithreaded performance).
As for PETSc's preference for processes over threads, we've found it to be every bit as fast as threads while offering more reliable placement/affinity and less opportunity for confusing user errors. OpenMP fork-join/barriers incur a similar latency cost to messaging, but accidental sharing is a concern and OpenMP applications are rarely written to minimize synchronization overhead as effectively as is common with MPI. PETSc can share memory between processes internally (e.g, MPI_Win_allocate_shared) to bypass the MPI stack within a node.
For dense matrices, I believe GEQP3 in LAPACK pivots so that the diagonal elements of R are decreasing, so we can just threshold and figure out when to cut things off. For sparse, the only code I've tried that's done this properly is SPQR with its rank-revealing features.
In truth, there may be a better way to do this, so I might as well ask: Is there a good way to find the generalized inverse of AA' where A is rank-deficient as well as short and fat?
As far as where they come from, it's related to finding minimum norm solutions to Ax=b even when A is rank-deficient. In my case, I know the solution exists for a given b, even though the solution may not exist in general.
Also, if your problem is a good fit for a method like this, it could be impetus to add it to PETSc. https://epubs.siam.org/doi/pdf/10.1137/120866580
But, yes, LSQR or more fitting LSMR solves a similar problem, but they're the iterative solver and I need the preconditioner, which I'm using the factorization for.
https://software.intel.com/content/www/us/en/develop/article...
I was multiplying and inverting sparse triangular matrices of size 650K x 650K with Matlab, on a laptop. Just amazing.
I'm baffled why there would be a problem with commercial users running a free software program like Julia or GNU Octave+SuiteSparse; that's Freedom 0. (And commercial /= proprietary, of course.)
That said, I believe it gets trickier once we start compiling the code. Say I want to develop a piece of software for my client and I don't want them to have the source, Octave doesn't really have a way to do this, but MATLAB does and since MATLAB has purchased all of the requisite licenses, we're good to go. Julia makes me more uncomfortable. We can make binaries with PackageCompiler.jl, but if we do, we should be subject to the provisions in the GPL. That's no different than any other piece of software, but Julia, Octave, and MATLAB all use these libraries and most people don't know that something like the chol command hooks into SuiteSparse in the backend.
SLIP_LU: GPL or LPGL
AMD: BSD3
BTF: LGPL
CAMD: BSD3
CCOLAMD: BSD3
CHOLMOD Check: LGPL
CHOLMOD Cholesky: LGPL
CHOLMOD Core: LGPL
CHOLMOD Demo: GPL
CHOLMOD Include: Various (mostly LGPL)
CHOLMOD MATLAB: GPL
CHOLMOD MatrixOps: GPL
CHOLMOD Modify: GPL
CHOLMOD Partition: LGPL
CHOLMOD Supernodal: GPL
CHOLMOD Tcov: GPL
CHOLMOD Valgrind: GPL
CHOLMOD COLAMD: BSD3
CPsarse: LGPL
CXSparse LGPL
GPUQREngine: GPL
KLU: LGPL
LDL: LGPL
MATLAB_Tools: BSD3
SuiteSparseCollection: GPL
SSMULT: GPL
RBio: GPL
SPQR: GPL
SuiteSparse_GPURuntime: GPL
UMFPACK: GPL
CSparse/ssget: BSD3
CXSparse/ssget: BSD3
GraphBLAS: Apache2
Mongoose: GPL
There's probably a bunch of mistakes in there, but that's what I found scraping things moderately quickly. Selfishly, I'd love SPQR to be LGPL, but everyone is free to choose a license as they see fit.It will probably be ported though, if there's a demand...
One question in case you or anyone else knows: What's the story behind AMD's apparent lack of math library development? Years ago, AMD and ACML as their high-performance BLAS competitor to MKL. Eventually, it hit end of life and became AOCL [2]. I've not tried it, but I'm sure it's fine. That said, Intel has done steady, consistent work on MKL and added a huge amount of really important functionality such as its sparse libraries. When it works, AMD has also benefited from this work as well, but I've also been surprised that they haven't made similar investments.
Also, in case anyone is wondering, ARM's competing library is called the Arm Performance Libraries. Not sure how well it works and it's only available under a commercial license. I just went to check and pricing is not immediately available. All that said, it looks to be dense BLAS/LAPACK along with FFT and no sparse.
[1] https://www.pugetsystems.com/labs/hpc/How-To-Use-MKL-with-AM...
It's ok. I did some experiments with transformer networks using libtorch. The numbers on a Ryzen 3700X were (sentences per second, 4 threads):
OpenBLAS: 83, BLIS: 69, AMD BLIS: 80, MKL: 119
On a Xeon Gold 6138:
OpenBLAS: 88, BLIS: 52, AMD BLIS: 59, MKL: 128
OpenBLAS was faster than AMD BLIS. But MKL beats everyone else by a wide margin because it has a special batched GEMM operation. Not only do they have very optimized kernels, they actively participate in the various ecosystems (such as PyTorch) and provide specialized implementations.
AMD is doing well with hardware, but it's surprising how much they drop the ball with ROCm and the CPU software ecosystem. (Of course, they are doing great work with open sourcing GPU drivers, AMDVLK, etc.)
https://developer.arm.com/tools-and-software/server-and-hpc/...
I don't see a story. AMD supports a proper libm for gcc and llvm, has its own libm, BLAD, LAPACK, ... at https://developer.amd.com/amd-aocl/
Just their rdrand intrinsic is broken on most ryzens if you didn't patch it. Fedora firmware doesn't patch it for you.
ACML was never competitive in my comparisons with Goto/OpenBLAS on a variety of opterons. It's been discarded, and AMD now use a somewhat enhanced version of BLIS.
BLIS is similar to, sometimes better than, ARMPL on aarch64, like thunderx2.
https://newsroom.intel.com/editorials/accelerating-foundry-i...
https://software.intel.com/content/www/us/en/develop/tools/m...
And, this question on Intel's own forums from 2016 at least suggests that there wasn't an MKL version for ARM in the time frame of the article you're linking to, either:
https://community.intel.com/t5/Intel-oneAPI-Math-Kernel-Libr...
So, from what I can tell, while Intel is an ARM licensee and made ARM CPUs in the past, they haven't made their own ARM CPUs for years and there's no sign they ever made MKL for any ARM platform. Never say never, but I think the OP is basically right -- there's not a lot of incentive for Intel to produce one.