.NET JIT and SIMD are getting married
blogs.msdn.com
blogs.msdn.com
[1] https://twitter.com/migueldeicaza/status/452099923157065728
Mono.SIMD has a few extra features that are missing from the current design, but they are not that important.
Microsoft has found a pleasant and easier to use API than we did. We have shared our feedback with them, and hopefully it will keep improving.
We're hoping to see this NuGet package supported on top of Mono, too. As we've done previously, we'll share our tests with Xamarin to ensure consistent behavior between implementations.
Oh well opening lines are what opening lines need to be. I'm just pedantic about stepping on solid ground.
25% isn't bad, but is no where near the 60% we were getting before. Especially since a 2 year wait used to be a 150% boost but is now a 56% one (with many analysts thinking that is generous).
[1] http://csgillespie.wordpress.com/2011/01/25/cpu-and-gpu-tren...
I wonder how it compares to a x86, but benchmarks are hard to come by.
That's false.
Moore's Law used to translate into increased clock rate. In consumer CPUs, this clock rate has stagnated at around 3 GHz ever since 2002 !!! This happens due to issues with power dissipation and increasing GHz count was a waste anyway, because performance increases from higher GHz isn't linear (i.e. a 25% increase in GHz count translates in less than 25% increase in performance and the difference in top of the line consumer CPUs running at 3 or 4 GHz and over-clocked prototypes running at 8 GHz is much smaller than you'd think ;))
Ever since then, single-threaded performance has been improving by design changes (e.g. instruction-level parallelism, caching, etc.), however this doesn't scale and isn't necessarily related to Moore's Law (i.e. the processor doesn't know the code's intention and so there are hard limits to how many smarts it can pull off).
https://en.wikipedia.org/wiki/Clock_rate#Historical_mileston...
See: http://preshing.com/20120208/a-look-back-at-single-threaded-...
SIMD can be a huge boost for numerical workloads; for example, SIMD auto-vectorization was just added to Julia, and people saw 8x boosts on some workloads (and in some cases another 4x from inlining improvements merged the next day).
Does the JVM have SIMD?
Having said this, each JVM vendor is free to make use of them if it sees fit to do so. They just don't offer a standard way to explore the SIMD instructions, or let the JIT decide how to compile the code.
Additionally, Java 9 will have integrated support for GPGPU via the Aparapi project for HSAIL support. Which you can already use anyway.
http://developer.amd.com/tools-and-sdks/heterogeneous-comput...
https://www.khronos.org/news/press/khronos-finalizes-opencl-...
"Host and device kernels can directly share complex, pointer-containing data structures such as trees and linked lists, providing significant programming flexibility and eliminating costly data transfers between host and devices."
OTOH, I find this specific C# implementation neat and portable. I like it a lot.
True, but if you're just trying to get some numerical speed out of some specific routines in a larger .NET application then having access to SSE like this could still be very helpful. Calling into native libraries from .NET isn't necessarily a performant option because of the cost of marshaling data back and forth between the managed and unmanaged memory spaces.
The end result can be pretty significant; from my own experience I'm usually pretty hard-pressed to come up with a C++ implementation that can beat the C# code it intends to replace outside of a microbenchmark. If the C# code now has the option of banging on SSE then I'm not sure it'll even be worth trying to trot out C++.
"Speed" is relative. If you mean throughput or scalability (2 different things), then the JVM or the CLR may be exactly what you want due to the ease with which one can juggle with multi-threading on multi-core processors.
Single-threaded performance is becoming less and less interesting and dealing with multithreading or with asynchronous I/O in lower-level languages, such as C/C++, is an extreme pain in the ass - because for example, C/C++ doesn't come with a memory model by itself (i.e. you can get fucked even when running with a different Glibc version) and the poster child for async I/O, libevent, has been plagued for years with concurrency issues, leading to a whole generation of insecure web servers.
The JVM, .NET and managed runtimes in general are great choices for going forward, because not only they come with memory model guarantees and a sane standard library - but having higher-level bytecode under the hood that can be generated at runtime means that either the runtime or the libraries running on top can repurpose that bytecode at runtime for optimal execution. Like for example, you could target the GPU when dealing with parallel collections.
It could just be that the current languages don't support it well.
Box2D for Java/C# is very useful (see Farseer in C# land), and can be more efficient than going through lots of interop calls. You can also hit DirectX directly using Sharp/SlimDX, which again, is also quite useful (if you want to write a game for DirectX, you are pretty much stuck in C++ unless you use these bindings).
I'm not hearing nearly as much about MonoGame as I am about Unity (although I am hearing some promising things.)
I don't mean to imply Unity is 1:1 API replacement for XNA - it isn't, MonoGame is, yes. But I'd argue Unity has replaced (and quite effectively at that) XNA/MonoGame as a strong game building target for the masses. And it's not even arguable that, for me, Unity has absolutely replaced XNA, and MonoGame at this point is background noise.
And don't get me wrong, I wish this weren't the case - because the more I've used Unity's APIs, the more I've come to hate them. I've an extreme dislike of the out of date and buggy compiler Unity uses for some of those platforms. But Unity's got a broader platform set than XNA and MonoGame combined, which unfortunately for me trumps the rest. I'd unscientifically wager it has far more libraries and middleware being written with it in mind.
Maybe that will change. I would welcome such change. There have already been changes, even. But I'm not holding my breath - even if it does, it's too little and far too late for my day job.
The tradeoff between safety/convenience/productivity vs. performance of C# vs C++ for example can't be motivated at the moment, and probably won't be with SIMD. Writing .NET code with deterministic performance is hard, so hard that C# is no more convenient than C++. I think any complex performance sensitive (Web browser, Game, SSL-Implementation...) is soon too complex to be done in C++ and too performance sensitive to be done in managed. Rust can't be released soon enough.
Then again the NHS here in the UK just rewrote a ton of stuff in RabbitMQ, Riak and Erlang from Oracle...
For performance reasons, we’ve defined those types as immutable value types.
Immutability strikes again.edit: On topic, this is really, really cool. I was playing around with SIMD intrinsics in C++ a few days ago (and realized that the compiler was in most cases generating equal or better code with the auto-vectorizer than me using intrinsics) so I have kind of a new-found interest in the topic.
First, making them immutable means you don't have to worry about memory barriers. That could be huge for data that's being shared by multiple threads. While these types logically have multiple elements, in reality they all fit within a single SSE register. Meaning the performance cost associated with having to worry about mutability could easily annihilate any potential performance boost you might get from being able to fiddle with the vector unit.
(Guess one-and-a-half is that, since these values are meant fit in a single CPU register, they're really more analogous to atomic types than they are to objects, anyway.)
Second guess is that it's more of a "pit of success" thing than a performance thing. Mutable value types in .NET are really problematic. I've seen so many bugs resulting from them that nowadays I consider them grounds for automatic rejection in code reviews.
Here's why. SIMD offers no new capabilities (1), only more speed, and not much more at that, maybe 4x if you're lucky. It's also hard to use: it requires unnatural data layouts, and lacks many operations (e.g. integer division). None of this is specific to the .NET implementation: it's just the nature of the beast.
So successfully exploiting SIMD is not easy, and requires thinking at the level where instruction counts matter. And because the amount of parallelism is so limited, high level languages (by which I include C!) can very easily blow away any gains with suboptimal codegen. Just a handful of additional instructions can ruin your performance (2).
Here's what will go wrong with an architecture-independent SIMD API:
1. Say you invoke an operation without an underlying native instruction. The compiler is forced to implement this by transferring the data to scalar registers, performing the operation, and then transferring the result back. Game over: this exercise is likely to eat up any performance benefit.
2. To avoid this, say you limit the API to some "common subset" of all extant SIMD ISAs. The problem is, many algorithms admit vectorization only through the exotic instructions, such as POPCNT on SSE4, or the legendary vec_perm on Altivec. If this instruction is not exposed, you can't vectorize the algorithm. Game over again.
That's why software that takes advantage of SIMD invariably has separate implementations for each supported ISA. .NET should have followed suit: expose an API for each ISA (or a mega-API that covers all ISAs), and then provide rich information about which operations are implemented efficiently, and which are not, to allow apps to choose an optimal implementation at runtime. This API would demo and market very poorly, but the engineers will love it, because it's the one that enables the most benefit from SIMD.
1: with rare exceptions, such as the new fused multiply-add support in x86
2: Several years back, VC++ generated an all-bits-1 register by loading it from memory, instead of issuing a pcmpeqd, which caused my vector implementation to underperform my scalar one. This is my fear for the .NET implementation.
That said you're right, usually the best performance can only be obtained by using really specific instructions. But in my experience, a decent performance increase can be obtained by using the generic vector extensions.
Moreover, if you can use the vector extensions for a large part of the code, that means you have to write a lot less platform-specific stuff. I.e.: you increase portability anyway, since now you only have to rewrite 5 out of 20 functions instead of 20/20. Even better, they allow one to write v3 = v1 + v2 instead of v3 = _mm_add_ps(v1,v2). The first one being clearer, more portable (will generate appropriate addps or equivalent NEON, ...) and plain nicer to read.
Your pcmpeqd example is a good example of an optimizer flaw. In my opinion this is orthogonal to whether or not to expose a specific or generic API. The compiler should've use the most efficient instruction for that simple idiom, period (without you telling it to use pcmpeqd). If we continue your line of reasoning, we're back to assembly for everything.
[1]: https://vec.io/posts/gcc-and-clang-vector-extensions (The vector extensions allow +,-,*,/,<,>,==,!= to be naturally used for SIMD types) [2]: https://github.com/rikusalminen/threedee-simd
Does anyone happen to know more about how/if shuffles are going to be exposed?
Prior art:
- LLVM's shufflevector intrinsic http://llvm.org/docs/LangRef.html#shufflevector-instruction (not meant for human consumption, but is an example of a multi-architecture backend with SIMD shuffle support)
- OpenCL's shuffle https://www.khronos.org/registry/cl/sdk/1.1/docs/man/xhtml/s... (Also note the `.s0123` "swizzle" syntax there.)
SSE has a bunch of other useful instructions like PMOVMSKB (useful for fetching the result of vectorized comparisons, yay!), then there are string instructions (sometimes also useful outside of string processing), etc.
New versions (AVX-512) will also have mask registers for masked operations.
If you really needed something like that, could you not use C++/CLI and expose those operations in your own unmanaged library? You will of course lose portability, but that seems like a possible work around.