Does the JVM have SIMD?
Does the JVM have SIMD?
Having said this, each JVM vendor is free to make use of them if it sees fit to do so. They just don't offer a standard way to explore the SIMD instructions, or let the JIT decide how to compile the code.
Additionally, Java 9 will have integrated support for GPGPU via the Aparapi project for HSAIL support. Which you can already use anyway.
http://developer.amd.com/tools-and-sdks/heterogeneous-comput...
https://www.khronos.org/news/press/khronos-finalizes-opencl-...
"Host and device kernels can directly share complex, pointer-containing data structures such as trees and linked lists, providing significant programming flexibility and eliminating costly data transfers between host and devices."
OTOH, I find this specific C# implementation neat and portable. I like it a lot.
True, but if you're just trying to get some numerical speed out of some specific routines in a larger .NET application then having access to SSE like this could still be very helpful. Calling into native libraries from .NET isn't necessarily a performant option because of the cost of marshaling data back and forth between the managed and unmanaged memory spaces.
The end result can be pretty significant; from my own experience I'm usually pretty hard-pressed to come up with a C++ implementation that can beat the C# code it intends to replace outside of a microbenchmark. If the C# code now has the option of banging on SSE then I'm not sure it'll even be worth trying to trot out C++.
"Speed" is relative. If you mean throughput or scalability (2 different things), then the JVM or the CLR may be exactly what you want due to the ease with which one can juggle with multi-threading on multi-core processors.
Single-threaded performance is becoming less and less interesting and dealing with multithreading or with asynchronous I/O in lower-level languages, such as C/C++, is an extreme pain in the ass - because for example, C/C++ doesn't come with a memory model by itself (i.e. you can get fucked even when running with a different Glibc version) and the poster child for async I/O, libevent, has been plagued for years with concurrency issues, leading to a whole generation of insecure web servers.
The JVM, .NET and managed runtimes in general are great choices for going forward, because not only they come with memory model guarantees and a sane standard library - but having higher-level bytecode under the hood that can be generated at runtime means that either the runtime or the libraries running on top can repurpose that bytecode at runtime for optimal execution. Like for example, you could target the GPU when dealing with parallel collections.