Sadly, (because KNL doesn't have nearly the same volume) until Skylake became a real thing, I had no reason to update this. I'm planning on dusting it off now!
FWIW we're currently using GCP, generally love it, and I'm looking forward to trying out Skylake...
For the numerical, AVX stuff -- it should be mostly automatic for you, if you're already using optimized libraries at the core -- MLK, BLAS, that kind of stuff. They'll be transparently upgraded for you -- ideally -- to take care of these things. They normally check what your CPU is at runtime, and pick the fastest implementation among a few different choices it has.
You will need toolchains to support this all, but for the most part that likely won't be a burden unless you want to get your hands dirty and start it yourself -- inevitably, this should all mostly be "pre-canned". Your optimized linear algebra, vector, and math libraries are what will mostly concern themselves with this, not you necessarily. In fact, several of the things already available can probably use these new extensions! I bet if you're using Intel MLK for example, it will probably "magically" get faster on these Skylake machines by using AVX512 automatically.
If you want to understand more: you can always go grab an SSE/AVX reference, check your /proc/cpuinfo, and write a few simple things on your own to get a feel. Your toolchain will definitely support it :)
(Aside, I think you mean MKL.)
There is a 200MHz clock reduction when running AVX512 instructions. If your code makes heavy use of AVX512 there is of course still a big net win, but I'm curious of the impact with more heterogeneous workloads. We have an app that is a mixture of scalar and vector code. Some, but not all, of the vector code would benefit from 512 bit vectors. But how much does the clock slowdown when running this code bleed over into running the other non-AVX512 code? I guess I'm asking how quickly it clocks down, and how quickly the full clock speed is restored. Worst case it seems you could be running full time at a 200MHz slowdown due to blocks AVX512 instructions scattered throughout the application. Is that a valid concern?
I'd have to look up the specifics; but does AVX512 simply slow the clock, or does it actually have some kind of limited number of hardware ports? I wonder if some clock slowdown would be very much of an issue, since clock-for-clock, you should see better performance on Skylake anyway.
Just curious, what kind of workloads do you think you're looking at here?
In my case, yes, Skylake would still be a win over older hardware, but the question is whether to use AVX512 or not. The workload is a real time animation system with a bunch of nodes in a graph that get evaluated in sequence. Some nodes would benefit from AVX512, but others would not. So the question is, if we vectorize those nodes that would benefit and get a speedup there, will the other unvectorized nodes now run slower as a result of the lower clock speed, canceling out the benefit.
It sounds like your case is a much better fit for AVX512. Out of curiosity, have you tried running on Xeon Phi, which also supports AVX512?
SGX (secure enclaves)
TSX (Hardware Transactional Memory)
SGX can store SSL keys in a hardware protected enclave that can't be accessed by hypervisors/AMT so that can be useful for security sensitive stuff.
Cloudflare probably could have used MPX to prevent their recent leak with minimal performance overhead.
TSX is cool.
Yes, it might have stopped the CloudFlare case, if they were willing to pay (I believe) for L4 over-read protection overheads (2x I believe, with a LOT of variance between GCC and Intel's compiler), and give up multithreading[1] in their application, and probably deal with other false positives and unsupported things.
SGX is completely neutered in Skylake and totally useless without getting your enclave signed by Intel, unless Google has struck a deal or something. You can mostly ignore it. It's good for preparation of your application against future processors where you'll have control over this, though, I guess...
TSX would be nice, yes. It's deeply annoying it's taken them so long to get right and that they've stratified that feature amongst CPUs -- is there any _real_ reason my Kaby Lake XPS13 can't support TSX? I doubt it other than "market segmentation makes us more money". I guess now I'm just ranting, though.
(As an example, my Xeon D-1540, a Broadwell-family chip, advertises the TSX bits in CPU feature flags, so it's not errata'd off there.)