Could you elaborate a bit on how JVM hinders reaching maximum FLOPS when multiplying sparse matrices?
I could think some examples where lacking SSE/AVX support would hinder it, but I don't see the connection with sparse matrices.
I could think some examples where lacking SSE/AVX support would hinder it, but I don't see the connection with sparse matrices.
A structural problem of JVM is that its runtime semantics is over-specified, there is very little room for the JIT to do its stuff. For example function arguments are evaluated right to left, there goes an opportunity for parallelism.
But how are sparse matrices then generally laid out? A naive approach would be some hash map, perhaps with some locality, in which I don't see JIT problems.