On a computer, such operations should almost never be implemented with scalar products, because the speed of a scalar product is limited by the latency of the FMA operation, not by its throughput. It is possible to accelerate the computation of a scalar product by ordering the elementary operations in a tree, but that complicates the program and it still does not reach the maximum speed of the hardware.
Except for vector-vector operations (i.e. for level 2 BLAS or greater levels), it is possible to change the order of the loops so that the scalar products are replaced by AXPY operations (whose result vectors should be kept in registers, to not be limited by the memory transfer throughput), which are not limited by the latency of the FMA, like the scalar products.
Except for vector-vector operations and matrix-vector operations (i.e. for level 3 BLAS or greater levels), it is possible to change the order of the loops so that the scalar products are replaced by rank-one matrix updates.
Most of the computation of a rank-one matrix update consists of the tensor product of 2 vectors.
Therefore, a matrix-matrix multiplication can be computed either by scalar products of the rows of the 1st matrix with the columns of the 2nd matrix, or by the tensor products of the columns of the 1st matrix with the rows of the 2nd matrix. The second method is much faster. The result of the tensor product must be kept in registers for the duration of the computation, so large matrices are partitioned in small blocks, i.e. submatrices, which are multiplied directly (an additional complication that increases the number of nested loops is that the small blocks that can be multiplied directly must be grouped in larger blocks that can be kept in the level-2 cache memory and reused for computations).
For matrix-matrix multiplications it is less important whether they are stored in column major order or in row major order, because when submatrices are brought into the L2 cache, the matrix elements will be gathered or scattered to be in the best order, before being loaded into registers repeatedly.
For matrix-vector multiplications, it is better to have the matrix in column major order, in order to compute the product by AXPY operations between columns and elements of the vector, and not by scalar products between rows and the vector (computing AXPY operations with shorter parts of a column at a time, so that the result vector can be kept in registers).