Comments indicate that's not a misprint. I'd love to hear the details of a 204 thousand percent improvement.
Comments indicate that's not a misprint. I'd love to hear the details of a 204 thousand percent improvement.
So I dunno about ZBrush. But Blender is a similar 3d program. There was a multithreading issue which would cause Blender to work for HOURS before actually completing some trivial tests.
Clearly, AMD Zen / Infinity Fabric has an edge case that older systems can handle just fine. Note that Blender uses Windows pthreads, which has an incredibly BAD implementation of spinlocks and mutexes.
Blender fixed this issue by making a customized implementation of spinlocks. I dunno what ZBrush did, but these kinds of architectural regressions happen every now and then.
Second: I'm not entirely sure if this applies to "modern" pthreads-win32, but the Blender project uses a relatively old build of pthreads-win32.
Third: my analysis is derived from the Blender diff here: https://dev-files.blender.org/file/data/rjmzozouj44bs6qicd3y...
I'm assuming that the Blender devs are smarter than me and have already tested this. :-)
------------
So going through the Blender source code we come across pthread_spin_lock as part of the diff. Considering the comment thread (as well as the diff) as adding Win32 #ifdefs and such, we can assume this was the cause of the performance issue.
---------
I assume... okay, so my logic isn't 100% tight, sorry :-(... that the implementation uses this pthreads-win32: http://sourceware.org/pthreads-win32/
Which brings us to here: ftp://sourceware.org/pub/pthreads-win32/sources/pthreads-w32-2-9-1-release/pthread_spin_lock.c
---------
Lets break it down why this is a bad implementation:
1. No "pause" instruction, causing latency issues when leaving the spinlock. See https://msdn.microsoft.com/en-us/library/windows/desktop/ms6... or https://software.intel.com/en-us/node/524249 for more information on why "pause" should be used in every spinlock. Strangely enough, the "pause" instruction is found in Linux implementations of pthreads (so the author for pthreads-w32 just... didn't look at the Linux implementation or something?)
2. PTW32_INTERLOCKED_COMPARE_EXCHANGE_LONG seems inefficient to me. I'd imagine that the usage should be using a swap instead (aka: https://docs.microsoft.com/en-us/windows-hardware/drivers/dd... ). Although I don't have any microbenchmark that "proves" this, I'd imagine that a swap is more efficient than a compare-and-swap.
3. A potential "slowpath" which includes a mutex-lock. Note that Windows Mutexes are incredibly heavy: including security features and even a "filesystem like" naming scheme. (See Windows Object Handles if you don't believe me). Windows Mutexes are very fully featured: recursive, multiple levels of security, and more.
This is a common problem in Linux system-coders -> Windows system coders. The Windows "Critical Section" is a more appropriate replacement for Linux pthread_mutexes.
I'd say Windows Mutexes might be more similar to a Linux filesystem-level semaphore. Linux semaphores provide those security features integrated with the filesystem.
The to Linux mutexes and spinlocks is something called a "Critical Section" in Windows land.
https://docs.microsoft.com/en-us/windows/desktop/Sync/critic...
Critical sections are Windows's lightweight yielding / task sleeping mechanism.
---------
Finally, Windows provides a spinlock: https://docs.microsoft.com/en-us/windows/desktop/api/synchap...
Which is probably is what pthreads-w32 should really be using. I mean, Linux devs basically just don't know what Windows offers, and the pthread-win32 library is a poor fit at the moment.
-----------------
Sorry if you were expecting more complete analysis and testing. Lol. Its mostly me looking at the thing and making a guess and reverse-engineering the patch-notes from the Blender discussion.
I did a brief search online, and Raymond Chen had this to say:
https://blogs.msdn.microsoft.com/oldnewthing/20160825-00/?p=...
> [snip] in real life, you should just use EnterCriticalSection because it has stuff like spin counts and lock convoy resistance.
Is there anything that pthread_mutex_lock does that EnterCriticalSection doesn't do? It seems like a good "translation" to me. I'd only use the WaitOnAddress thing if pthread_mutex_lock had some features that a raw EnterCriticalSection wouldn't do.
In addition, pthread mutexes can be statically initialised which requires some kind of initialise-on-first-use for those if you implement them using Critical Section objects.
I believe WaitOnAddress(), rather than being a "competitor" for futex was actually added (or at least, the underlying infrastructure was) in order to be able to implement futex() in the WSL.
I'm also curious what the exact issue was, but given the scale of improvement, I wouldn't be surprised if it is some sort of cache contention between threads/processors. A cache miss in the inner loop of an algorithm can be catastrophic, and apparently Ryzen has pretty different cache behavior versus other CPUs:
https://www.reddit.com/r/Amd/comments/5x7oaq/ryzens_memory_l...
When I worked at a visual effect software house, AMD would send us all sorts of things. The main problem was thier linux drivers were terrible.
A lot of the time thier hardware was faster at certain things. At one point (this is about the time of the quadro 2/3/4/5/6/000) the AMD firepro was 8 times as fast. However their driver support was terrible, and the time to fix was 6months+
There was a reason why the mac pro had two AMD GPUs in them, they were at the time light years faster.(and cheaper.)
I don't know actual details of that, but words on the internet said that it is just a bug fixed.
As such, we programmers find GREAT interest in these sorts of bug fixes. Every bug is an opportunity for a programmer to learn more about computers!
Most bugs I've fixed have taught me relatively little about computers but lots about myself.