Intel Compiler Intrinsics Guide
software.intel.com
software.intel.com
Otherwise, this is one of the best cpu documentation tools I've used.
As a side request, it would be amazing if the intel documentation on non-temporal memory operations and sfence+lfence had more specifics about how they interact with the rest of the cache hierarchy and the load/store subsystems.
Since version 3.3.16, including the current 3.4, the guide no longer contains this info.
Why have you removed that?
Nice visual design.
Searching for "dot product" does not work, but searching for "dot-product" does work.
Also, every character I type in the search box ends up as a new entry in my browser history ...
How does hardware video decoding and encoding work?
Does the fact that modern processors have on-chip GPUs with acceleration mean the instructions are any different?
(Graphics has long been a "???" of mine, but I don't want to get into generalizations in this particular thread. Hardware video {de,en}coding seems mildly relevant though.)
GPUs themselves are mostly SIMD vector processors for their "shader cores" with a bunch of custom fixed-function hardware for the more special blocks in the pipeline. This is a completely separate unit from the CPU. There was an attempt from Intel known as "Larabee" to try and build a GPU style pipeline on top of an expanded Intel CPU. The consumer product was canned but the expanded CPU went on to be known as Intel's Xeon Phi line.
For example, _mm_add_pi8() seems to be identical to _m_paddb(). Same for _m_empty() and _mm_empty().
Are there any subtleties I'm missing?
I think the instructions which didn’t need to change just got a new intrinsic alias so that each instruction set is self-contained, i.e. when working with SSE, you should only need to look at the SSE docs, not also know the MMX docs already.
A version of the godbolt compiler explorer that included execution would be cool.
Hm, it could actually make for a fun hack.
Adds to todo list
ARM instructions are always 32 bit; this includes NEON. There’re signs ARM developer indeed were trying to minimize their count (e.g. there’s no right shift neon instruction, instead left shift is used with negative shift value).
Thumb instructions can be 16 bit, but still this is way simpler than x86, where a single instruction can be between 8 and 120 bits.
CISC vs RISC for modern CPUs (last 20 or 30 years) doesn't mean anything. Any modern ARM or Intel is partially RISC and a CISC at the same time.
Think this way:
1) modern x86 micro-ops can be viewed as RISC. So x86 is RISC?
2) ARM1 (from 1985) had an micro-ops, that means user visible opcodes are not reduced enough. So ARM it CISC? http://www.righto.com/2016/02/reverse-engineering-arm1-proce... Not talking even about modern 64-bit arm cores.
Also, ARM says their CPU designs are based on RISC principles: https://developer.arm.com/products/architecture/cpu-architec...
You basically are writing assembly with these intrinsics.
I'm guessing these intrinics map 1:1 to some AMD equivalent. Oh, looky here: https://msdn.microsoft.com/en-us/library/hh977022.aspx