“as_strided” and “sum” are all you need
jott.live
jott.live
https://ajcr.net/Basic-guide-to-einsum/
> The einsum function is one of NumPy’s jewels. It can often outperform familiar array functions in terms of speed and memory efficiency, thanks to its expressive power and smart loops
it seems like einsum has been ported to torch & tensorflow
https://numpy.org/doc/stable/reference/generated/numpy.einsu... ; https://pytorch.org/docs/stable/generated/torch.einsum.html ; https://www.tensorflow.org/api_docs/python/tf/einsum
1. Einsum doesn’t AFAIK allow custom strides. So Einsum does only the mapreduce part (and that too only for multiplication & summation combo). To get a taste for the absolute brilliance of a slightly more generic interface, check out Tullio in Julia.
2. Tensordot etc use BLAS under the hood IIRC and can therefore parallelize automatically. The Einsum backend is good old fashioned single threaded C code, so it’s fast (and inflexible) but single core.
I don't think there exists a cuda-accelerated einsum, so the equivalent matmul is always faster? With cuda, it might be implemented by "exploding" the inputs to matrices, which means it can quickly run out of memory?
This doesn't work as an explanation without defining what a "stride" is. You can't explain addition just by showing some examples of "before addition" and "after addition".
Here's what I found at https://ipython-books.github.io/46-using-stride-tricks-with-...:
> Strides describe, in any dimension, how many bytes we need to jump over in the data buffer to go from one item to the next.
So the stride in a 1-D array of float64 is 8 because you need to go 8 bytes forward to go from one item to the next. Change it to 32 and you have a way to access every fourth item without defining a new array.
Try soap instead.