Yeppp: A High-Performance SIMD-Optimized Math Library for X86, ARM, and MIPS
yeppp.info
yeppp.info
Note to mods: it would be fair to add (2013) to the link, as the site wasn't updated since Yeppp 1.0.0 release
Looks like a great approach!
Do you have thoughts on how this could be extended to work with Xeon Phi and graphics coprocessors? The particular grail I'm searching for would do what you have done, but also allow some operations to be offloaded.
Also, I see you have some preliminary R bindings: https://bitbucket.org/MDukhan/yeppp/src/7830144789416f9fbed3...
Are these thought to be working?
The R bindings are working and do provide speedup, but they are in no way of release quality. The officially supported bindings (.Net, JVM, and FORTRAN) are auto-generated. Additionally, Julia team maintains Yeppp.jl - bindings for Julia (https://github.com/JuliaLang/Yeppp.jl)
EDIT: I looked at the functions listed for multiplication. Although it has a 64-bit x 64-bit multiply, the result is not in the internal 128-bit format. Also there was a warning that 64bx64b multiply was not yet optimized.
Still very cool though.
FWIW, you don't tend to see a significant gain from writing those individual types as SIMD. The gain is from operations over large arrays of numbers. E.g. each vector has only one component type in it (e.g. you have vectors of [x0, x1, x2, x3] and [y0, y1, y2, y3] instead of [x0, y0, z0, w0] etc).
Here's a set of slides that elaborate further on what I mean and explain why: https://deplinenoise.files.wordpress.com/2015/03/gdc2015_afr...
Writing asm or using intrinsics seems like such a pain and non-portable.
But that doesn't explain the lack of language support for 3- and 4-element vectors (or even 2d). I think the GCC intrinsics are nice, they allow you to pass by value and are cross-platform with x86, ARM NEON, and AltiVec.
So many people think SIMD is somehow about parallel computation, but to me the obvious use is small vector math with basic vector types and standard operators.
The fact that you think that indicates that you have no experience with SIMD programming or instruction sets. Using SIMD for small vector math is at best a wash, and typically a poor idea for performance. You only get significant performance benefit out of SIMD instructions when you can use vectorize an entire inner loop (moving between scalar and vector ALUs is fairly expensive on some microarchitectures), and when data arrangement and movement is arranged to maximize throughput (think SoA instead of AoS). Using SIMD for short vector math is generally moving in the opposite direction of those goals, resulting neutral to negative performance delta compared to scalar code.
Basically, instead of trying to gain a bit performance out of one [x, y, z, w], you group your data so you have to process all of them at once, and then you load them as [x0, x1, x2, x3], [y0, y1, y2, y3], [z0, z1, z2, z3], [w0, w1, w2, w3], and then do 4 (or 8, or 16) steps of the loop at a time.
This ends up having roughly a 4 (or 8, or 16) times speed up. More if your data wasn't already in that format, which is inherently cache friendly.
I linked this elsewhere in these comments, but here's a good overview of SIMD techniques (as well as what not to do, which is basically what you had suggested) https://deplinenoise.files.wordpress.com/2015/03/gdc2015_afr...
I see this line of code, and I want to write the comment as code, rather than the code:
__m128 dw = _mm_mul_ps(aw, bw); // dw = aw * bwI just want to have a vector X = [1,2,3,4] and Y = [5,6,7,8] and be able to write A = X+Y. Or for a constant k, write A += k*X. These are extremely fast when using vector operations on intrinsics (one instruction). 2,3, and 4-element vectors are so common, it makes perfect sense to support them directly in a language.
Yes you can do much better by hand-writing assembly, but right now that's rarely done. So compilers try to vectorize scalar code we give it in languages like C, which is much harder than optimizing code that's already written in vectors like int32x4 or float64x2.