It is not technically legal to use SIMD on most floating-point loops, including the inner product in matrix multiplication, because rounding errors are not commutative. C compilers don't vectorize such loops either unless you pass the -ffast-math flag. I'm sure the JIT compiler of JVM has a similar option.
(read the article I linked below, it's all explained there)