Exploring .NET Core platform intrinsics: Part 4 – Alignment and pipelining
mijailovic.net
mijailovic.net
sum = Avx2.Add(block0, sum);
sum = Avx2.Add(block1, sum);
sum = Avx2.Add(block2, sum);
sum = Avx2.Add(block3, sum);
have all a serializing dependency on sum variable. But (integer) addition is associative and commutative, so you could sum it in a tree-like manner, ending up only with a a single serializing dependency: sum01 = Avx.Add(block0, block1);
sum23 = Avx.Add(block2, block3); // These two run in parallel
sum = Avx.Add(sum, sum01); // sum01 hopefully ready; parallel with sum23
sum = Avx.Add(sum, sum23); // sum23 hopefully ready
Where only the last line serializes with the previous one. Maybe the HW is smart enough to rename the registers and do the same thing internally, but it'd be interesting to benchmark it.For what it's worth, the vmovdqa only has a 4-wide issue width if it is moving between registers, the memory load has a 2-wide issue width. Floating point adders themselves only have a 1-2 wide issue widths depending on your hardware so it doesn't really matter.
It's remarkable how fast the language and runtime have evolved for performance. It wasn't that long ago that I was manually inlining Vector3 operators to try to get a few extra cycles out of XNA on the Xbox360.
Hence killing XNA when they took over Windows 8 development, WinRT and such.
It took all the reorganizations and change of politics, for the .NET Runtime finally start getting some additional love regarding performance.
The past tense structure makes it sound like progress has been made on this front while it’s still the same problem presently. It’s just that those tools in particular have been deprecated (and not replaced)
That was the reason why the XBox 360 runtime was bad, the remaining of my comment refers to the standard .NET Framework.