Update: I ran a quick benchmark (large matrix multiplication) between GNU APL, J, and NumPy (standard pre-compiled linux amd64 packages, all running single-core.) Here are the results.
GNU APL (1.7)
⍴+.×⌿?2 3000 3000⍴1e10
- size 3000: 65s, 9GiB RSS
(crash on bigger sizes)
J (8.07) $(+/ .*)/?2 3000 3000$1e10
- size 3000: 2s, 0.2GiB RSS
- size 10000: 60s, 3.7GiB RSS
(crash on bigger sizes)
NumPy (1.13.3, blas/lapack 3.7.1) import numpy
a=numpy.random.randint(0, 1e10, (2,3000,3000))
print((a[0,] @ a[1,]).shape)
- size 3000: 25s, 0.2GiB RSS
- size 10000: (>15m, I killed it) 2.3GiB RSS
Conclusions:J's implementation is surprisingly performant! Easily beating NumPy on speed alone by a factor of 10 or more! (And I thought blas/lapack were already heavily optimized libraries!) J's memory usage is comparable to that of NumPy. GNU APL had the worst memory and cpu profile of them all.