After all most tensors are implemented as 1D arrays with strides for each dimension with the innermost dimension always contiguous.
After all most tensors are implemented as 1D arrays with strides for each dimension with the innermost dimension always contiguous.
If the OP cares about implementation details about how an API like PyTorch is made, I think the MiniTorch 'book' is a pretty good intro:
I think accelerators mostly do computations over small matrices and not vectors, so it is "matricized" code, but I am not an expert in this area.
Sure! Let me break it down for you:
When we talk about "tensor ops," we're referring to operations performed on tensors, which are multidimensional arrays of numbers commonly used in mathematics and computer science.
Now, these tensor operations are typically implemented using a technique called "vectorizing code." In simple terms, vectorizing code means performing operations on entire arrays of data instead of looping through each element one by one. It's like doing multiple calculations at once, which can be more efficient and faster.
Tensors are usually represented as 1D arrays, meaning all the elements are arranged in a single line. Each dimension of the tensor has a concept called "stride," which represents how many elements we need to skip to move to the next element in that dimension. This helps us efficiently access and manipulate the data in the tensor.
Additionally, the innermost dimension of a tensor is always "contiguous," which means the elements are stored sequentially without any gaps. This arrangement also aids in efficient processing of the tensor data.
So, in summary, tensor ops involve performing operations on multidimensional arrays, and we use vectorized code to do these operations more efficiently. Tensors are represented as 1D arrays with strides for each dimension, and the innermost dimension is always stored sequentially without any gaps.