Python 4.216 GFLOPS
Naive: 6.400 GFLOPS 1.52x faster than Python
Vectorized: 22.232 GFLOPS 5.27x faster than Python
Parallelized: 52.591 GFLOPS 12.47x faster than Python
Tiled: 60.888 GFLOPS 14.44x faster than Python
Unrolled: 62.514 GFLOPS 14.83x faster than Python
Accumulated: 506.209 GFLOPS 120.07x faster than Python* For context I do have done some experience experimenting on the gcc/intel compiler options that are available for linear algebra, and even outside of BLAS, compiling with -o3 -ffast-math -funroll-loops etc does a lot of that, and for simple loops as in matrix vector multiplication, compilers can easily vectorize. I'm very curious if there is something I don't know about that will result in a speedup. See e.g. https://gist.github.com/rbitr/3b86154f78a0f0832e8bd171615236... for some basic playing around
I'm not sure where/how they'd be squeezing out more performance unless its better compilation/compatibility with Apple Silicon intrinsics.
Edit: ..Is Mojo using more than 1 core? I'm not sure I understand their syntax and if they are parallel constructs.
Edit2: Yeah Mojo seems to be parallelizing, so the comparison really isn't fair. The np.config posted elsewhere shows that OpenBLAS is only compiled with MAX_THREADS=3 support, and its not clear what their OPENBLAS_NUM_THREADS/OPENMP_NUM_THREADS was set to at runtime.
Python 119.189 GFLOPS
Naive: 6.275 GFLOPS 0.05x faster than Python
Vectorized: 22.259 GFLOPS 0.19x faster than Python
Parallelized: 50.258 GFLOPS 0.42x faster than Python
Tiled: 59.692 GFLOPS 0.50x faster than Python
Unrolled: 62.165 GFLOPS 0.52x faster than Python
Accumulated: 565.240 GFLOPS 4.74x faster than Python
np.__config__: Build Dependencies:
blas:
detection method: pkgconfig
found: true
include directory: /opt/arm64-builds/include
lib directory: /opt/arm64-builds/lib
name: openblas64
openblas configuration: USE_64BITINT=1 DYNAMIC_ARCH=1 DYNAMIC_OLDER= NO_CBLAS=
NO_LAPACK= NO_LAPACKE= NO_AFFINITY=1 USE_OPENMP= SANDYBRIDGE MAX_THREADS=3
pc file directory: /usr/local/lib/pkgconfig
version: 0.3.23.dev
lapack:
detection method: internal
found: true
include directory: unknown
lib directory: unknown
name: dep4364960240
openblas configuration: unknown
pc file directory: unknown
version: 1.26.1
Compilers:
c:
commands: cc
linker: ld64
name: clang
version: 14.0.0
c++:
commands: c++
linker: ld64
name: clang
version: 14.0.0
cython:
commands: cython
linker: cython
name: cython
version: 3.0.3
Machine Information:
build:
cpu: aarch64
endian: little
family: aarch64
system: darwin
host:
cpu: aarch64
endian: little
family: aarch64
system: darwin
Python Information:
path: /private/var/folders/76/zy5ktkns50v6gt5g8r0sf6sc0000gn/T/cibw-run-27utctq_/cp310-macosx_arm64/build/venv/bin/python
version: '3.10'
SIMD Extensions:
baseline:
- NEON
- NEON_FP16
- NEON_VFPV4
- ASIMD
found:
- ASIMDHP
not found:
- ASIMDFHMBecause it's slow as dirt right? Isn't the point that they are trying to make is that one could with Mojo?
[1] https://github.com/modularml/mojo/blob/5ce18c47a27c0c4123de1...
In a single-threaded comparison to numpy (or just measuring total throughput -- many applications have lots of slightly smaller matmuls they can do which make it trivially to parallelize without having to parallelize each matmul, and throughput increases slightly when you do so) though, the details start to matter. Numpy is bad with small dimensions (hasn't optimized for them at all really, and overhead moving data from a Python context to a Numpy context starts to dominate), and performance can vary 10-50x just based on whether you've set up an optimized BLAS library for it to link to or not. Mojo side-steps some of that because it provides the fast primitives you need in the language itself and doesn't present the opportunity to execute more slowly. Single-core Mojo shouldn't be meaningfully faster than a properly installed single-core Numpy on large matmuls, and the given implementation should be meaningfully slower on large enough problems.
I don't really care for the benchmark though. It's potentially okay at showing how easy it can be to write fast code, but it comes across as being presented to show how much faster mojo is than Python. That latter is misleading for at least a couple reasons:
- By some magic, my $300 old dev laptop (swift 3) is 6x faster than their brand new m2 pro max with a vanilla Python triple for loop. Is Mojo adding some overhead as it runs that benchmark? Is something wrong with their Python installation?
- Many of the optimizations they applied in Mojo apply just as well in vanilla Python. Tiling, parallelization, ... Some of those have a higher ratio of improvement in Python than Mojo (depending on some fiddly GC details) because their purpose is to cut back on cache/page/... misses enough to make the problem compute-bound rather than IO-bound, and the Python representation takes enough extra space that the benefits accumulate faster. The serialization overhead for stdlib parallelization is dwarfed by the matmul cost, so it doesn't wind up mattering much that you have slow copies all over the place, and you really do find yourself bound by the interpreters rate of interpreting.
Like, it's still faster than Vanilla Python by a lot, and it's neat that the code is so easy to write, but 90k isn't the speedup I'd headline with.
To be fair, I think their point is just that you can write fast code easily in Mojo, and matmul is something easy to understand, so it makes a good case study. The optimization primitives are fairly intuitive, so presumably you should be able to apply the same naive approach (code it, slap on optimization primitives) and get decent speedups in less well studied domains.