The source is in the text, but conveniently accessible here: https://github.com/uncomplicate/neanderthal/blob/master/exam...
The source is in the text, but conveniently accessible here: https://github.com/uncomplicate/neanderthal/blob/master/exam...
For completeness sake, could you also provide some information about your machine specs and operating system (and if on linux your glibc and kernel version)?
ND4J returns the `mmuli` result in F order, and DL$J is aware of that. We design our algorithms around this. It is possible to get the `mmuli` result in C order, but it'll cause an additional conversion operation call, which is expensive. That's exactly what Dragan did. So on one hand - he's right: in the CCC case, Neanderthal is faster, because he doesn't have to do the F->C conversion of the result.
We make mmuli always return F because of Cuda. The cuBLAS implementation has no option for C ordered output.
`mmuli` CCC always introduces a dup for us. We expect the gemm result to always be F, so if result is not F, each operation allocates tempResult array, F ordered, and does result.assign(tempResult).
Thanks for the explanation. I used the code that ND4J guys provided, assuming that they'd do the right thing. Also read my response to crockpotveggies: provide the ND4J code that you think is optimal, and ND4J and Neanderthal results on your machine, and I'll be happy to write a follow up.
transpose is the only option.
cublasStatus_t cublasSgemm(cublasHandle_t handle, cublasOperation_t transa, cublasOperation_t transb, int m, int n, int k, const float alpha, const float A, int lda, const float B, int ldb, const float beta, float *C, int ldc)
https://neanderthal.uncomplicate.org/articles/getting_starte...
And just to be sure, choose option 2 under that heading (add the folder with appropriate dlls to PATH)