An Intel Programmer Jumps over Wall: First Impressions of ARM SIMD Programming
branchfree.org
branchfree.org
[ Note the title of the article was "Jumps over The Wall", in keeping with the dish. ]
"aXX optimization guide"
a57: http://infocenter.arm.com/help/topic/com.arm.doc.uan0015b/Co...
a72: http://infocenter.arm.com/help/topic/com.arm.doc.uan0016a/co...
Notably not public is the a53 version of this document. I did some work getting ORB_SLAM to run on the raspberry pi a couple years ago. My conclusion was that the instruction timings for NEON instructions are more or less the same as the a57, but just remember there are only two instructions issued per cycle and everything is in order.
Edit: replace google with ddg
NVidia publishes throughput metrics in their PTX page: https://docs.nvidia.com/cuda/cuda-c-programming-guide/index....
AMD's information is harder to find, but it does exist: http://clrx.nativeboinc.org/wiki2/wiki/wiki/GcnTimings
--------
I haven't been doing too much SIMD programming myself, but I know the general intrinsics. I think ARM's biggest problem is documentation. Intel's intrinsics page is great: https://software.intel.com/sites/landingpage/IntrinsicsGuide...
ARM's intrinsics page is... not as good. But I guess still workable. As you've noted, ARM's intrinsics documentation is missing latency / throughput metrics. https://developer.arm.com/technologies/neon/intrinsics
POWER9 by the way has latency/throughput metrics published. I'm most happy with POWER9 as an alternative CPU, too bad its so expensive though.
----------
Intel + Agner Fog is definitely the best documentation for how things work and optimization information. NVidia CUDA and POWER9 seem to be #2 and #3 with documentation.
Since GPUs are available for cheap these days (under $300 and they plug into any motherboard), GPUs seem to be the best alternative computer these days to learn. Since they're decently documented, its good to explore IMO.
I keep meaning to go back and do some GPGPU again. I am very fond of architectures that actually exist. GPU is a very different bet than SIMD but a lot of fun.
Incidentally PTX is a fiction - more of an IR than something with true throughput and latency numbers. I don't know if anyone programs GPUs directly but that would be interesting.
PTX is a fiction, but its very close to the underlying "SASS" assembly language of NVidia GPUs. There was some research done into this field here: https://arxiv.org/abs/1804.06826
They reverse engineered the SASS assembly language, all the op-codes match up to PTX opcodes. The main difference is that SASS includes control information, such as stall cycles, write or read barrier information.
In effect, it seems like NVidia's compiler figures out stalls and dependencies at compile-time. Memory-dependencies are written to each SASS assembly instruction in Volta, in a format that has changed differently each generation.
The typical assembly programmer doesn't care about these details (CPUs typically handle that logic automatically), so it makes sense to ignore them through the higher-level PTX Assembly language.
It seems like PTX is sufficient for understanding the general execution of an NVidia GPU. There's probably no reason to write the underlying SASS code by hand, especially because PTX matches up to the SASS opcodes. Furthermore, it is clear that NVidia is tweaking the details of those memory-barrier and dependency information, and doesn't want programmers to write against that abstraction level.
In the paper, we do also include measurements against 7 other libraries including jsoncpp. It is an order of magnitude slower than we are on the measurements we took. I am not familiar enough with it to know whether that's meaningful; it may have a lot of other attributes that make it more useful for what you're doing than what we have. We're pretty barebones, especially at this stage.
For example, there are some nice possibilities with the TBL instruction (or PSHUFB, or VERMB, etc) for doing character class membership tests. Which version you use would have a lot to do with whether you think TBL is going to issue 2/cycle, 1/cycle or 0.5/cycle (just fr'instance) - you might design the algorithm quite differently. It's not just a case of picking the one design choice and then having the compiler schedule that - if you're not thinking about this during design, you're not going to be remotely near peak performance.
(edit - also worth noting that this question seems to refer to my old job; I don't currently work on Hyperscan, although I may go back to regex one day. I'm sympathetic with the line of questioning, though, so please stop downvoting the parent)
Used to be Scalable VE
I keep meaning to write more about this. I have found 3 categories of SIMD use in my own work:
1) Doing one thing a gazillion times (e.g. conventional SIMD). This works well on vector machines, of course.
2) Using SIMD registers to do "more stuff". So in a couple string matchers and regex matchers I've designed, you are really use using a SIMD register because it's bigger than a GPR register (duh). But this might be because you want to simulate a 512-bit NFA rather than a 64-bit NFA.
3) Using SIMD operations to do weird, irregular stuff where what you're really getting is a substitute for branchy code. I blogged about an example of this called "SMH" https://branchfree.org/2018/05/30/smh-the-swiss-army-chainsa...
I suppose it's hard to know until we can actually write code for it, but I haven't imagined a scenario where the RV model is significantly worse than packed SIMD, other than a tiny bit of vector configuration bookkeeping. I think that's a worthwhile tradeoff for getting simple, portable, fast code for elementwise operations and implementation flexibility.
I'd love to be convinced otherwise, though.
#pragma vectorize_me_to_death
for (int i = 0; i < N; i++) {
// ...
}
But there is also a second kind of vector opportunity: small vector opportunities. SLP vectorization is the ur-example here: you scan a block of code for operations that happen to be doing the same operation on different values and make vector code out of it. For this kind of vector code, there is a lot more focus on horizontal and shuffling code than the wide kind of vector code.I haven't looked at the ARM SVE or RISC-V vector ISAs in detail, but I imagine that they don't support the latter kind of vectorization very well.
A "tiny bit of vector configuration bookkeeping" has my hackles going up. I'm used to SIMD operations where it's quite usual that you have latency = 1 and reciprocal throughput = 3. This is a narrow path to walk and not one where "tiny bit" of extra overheads will be welcome. I guess we'll see - I would like RISC-V to succeed, but worry that all these resizable models will be quite slow for the codes I write now.
The saving sequence would be something like this (after setting up a frame pointer if needed): "vsetvl t0, x0, e8, m8; sub sp, sp, t0; vse.v v16, (sp)".