Tensors are really different mathematical objects with a far more rich structure than those used in the Deep Learning context.
https://en.wikipedia.org/wiki/Vector_(mathematics_and_physic...
In linear algebra you should learn this as soon as you hear about it. Is the english Wikipedia page correct in that you always use nabla for the gradient? https://en.wikipedia.org/wiki/Gradient#Generalizations
Because `grad f = ∇f` holds in scalar fields f only (not in vector or tensor fields (tensors that aren't scalars)).
In the ML context it generally doesn't matter because you don't do these transformations (and more generally there is no metric for your space). So you can just treat a gradient as an array of numbers. But in a physical context the distinction starts to matter, at least on a manifold with curvature.
The one place it does matter in ML (that I'm aware of) is information geometry, where you try to do a transformation from the normal coordinate space to a more "natural" coordinate system that is based on the Fisher information, and this introduces curvature into model.
Would indeed be cool to have true tensor processing units :-)
PKD level of cool as you could argue those chips directly process spacetime chunks
Just calling it an array might be underselling it, at this point. Perhaps a tensor in the context of CompSci is just a particular type of array with certain expected properties.
I like the NDArray terminology used by numpy. I think it gets the point across more clearly than “tensor”, and conceptually they’re pretty much equivalent
After all most tensors are implemented as 1D arrays with strides for each dimension with the innermost dimension always contiguous.
Sure! Let me break it down for you:
When we talk about "tensor ops," we're referring to operations performed on tensors, which are multidimensional arrays of numbers commonly used in mathematics and computer science.
Now, these tensor operations are typically implemented using a technique called "vectorizing code." In simple terms, vectorizing code means performing operations on entire arrays of data instead of looping through each element one by one. It's like doing multiple calculations at once, which can be more efficient and faster.
Tensors are usually represented as 1D arrays, meaning all the elements are arranged in a single line. Each dimension of the tensor has a concept called "stride," which represents how many elements we need to skip to move to the next element in that dimension. This helps us efficiently access and manipulate the data in the tensor.
Additionally, the innermost dimension of a tensor is always "contiguous," which means the elements are stored sequentially without any gaps. This arrangement also aids in efficient processing of the tensor data.
So, in summary, tensor ops involve performing operations on multidimensional arrays, and we use vectorized code to do these operations more efficiently. Tensors are represented as 1D arrays with strides for each dimension, and the innermost dimension is always stored sequentially without any gaps.
If the OP cares about implementation details about how an API like PyTorch is made, I think the MiniTorch 'book' is a pretty good intro:
I think accelerators mostly do computations over small matrices and not vectors, so it is "matricized" code, but I am not an expert in this area.
A good basic thing to know is a tensor is a single dimensional array under the hood. At least as far as I know in my mental model - I have not yet looked at the code :-)
Therefore "converting" a 2 * 5 * 10 tensor into a 2 * 50 one is an immediate operation, it doesn't have to scan. Just change some metadata.
A 2-dimensional array A will, in the object program, be stored sequentially in
the order Ai,i, A2,i Am l , A| i2 , A2f2 Am, 2> , Am,„. Thus
it is stored “columnwise”, with the first of its subscripts varying most
rapidly, and the last varying least rapidly. The same is true of 3-dimensional arrays.
1 -dimensional arrays are of course simply stored sequentially. All arrays are stored backwards in storage; i.e. the above sequence is in the order of decreasing absolute location.
https://ia800807.us.archive.org/0/items/history-of-fortran/I...This is overly simplified though. Things get different when we start talking about {S,M}I{S,M}D (page addresses SIMD) or GPUs. Parallel computing is a whole other game, and CUDA takes it to the next level (lots of people can write CUDA kernels, not a lot of people can write GOOD CUDA kernels (I'm definitely not a pro, but know some)). Parallelism adds another level of complexity due to being able to access different memory locations simultaneously. If you think concurrent optimization is magic, parallel optimization is wizardry (I still believe this is true)