Why Spinlocks Are Bad On iOS
engineering.postmates.com
engineering.postmates.com
Is this true? OSSpinLock (at least on OS X) calls thread_switch for 1ms every 1024 iterations (which as far as I understand is basically sleeping[1]). Presumably it's not runnable while it's sleeping?
After 100 * 1024 iterations it will also ask to be throttled to the lowest priority for 1ms, again with thread_switch.
[1] http://web.mit.edu/darwin/src/modules/xnu/osfmk/man/thread_s...
AKA an adaptive mutex.
Assuming you mean something like http://stackoverflow.com/a/25168942/582, you basically need a full mutex anyway, so I'm not sure why you wouldn't just use a mutex directly. I don't think you can really implement this adaptive mutex in userland because you can't just use a spinlock + a mutex (you'd have to lock the mutex every time anyway, which defeats the point of having the spinlock) and syscalls are not considered public API (the kernel is allowed to change syscalls between versions, libc/libsystem is the only public API for them).
FWIW pthread mutexes were sped up 2-2.5x in "new OSs" to compensate for spinlocks being illegal (https://twitter.com/catfish_man/status/676852111527706624).
I don't understand. Could you explain this in more detail?
In Linux, adaptive mutexes are implemented entirely in userspace with futex as the primitive syscall. This is because atomic CAS (cmpxchg, ldrex/strex, etc) suffices to ensure synchronization during the spinning.
FWIW, libdispatch itself definitely uses Mach SPI. For example, skimming through the source, it calls thread_switch() using a non-public option (there's 3 non-public options to thread_switch now, two of which are intended for dealing with the QOS issues that my article talked about; the private os_lock_handoff_s that the objc runtime uses internally uses one called SWITCH_OPTION_OSLOCK_WAIT and libdispatch appears to use the other, SWITCH_OPTION_OSLOCK_DEPRESS, as well as the third option which was added just for libdispatch called SWITCH_OOPTION_DISPATCH_CONTENTION).
In general, Apple has never encouraged anyone to use spinlocks, and as I mentioned at the end, as long as your critical section is really tiny and CPU-bound you're unlikely to actually have a problem in practice. Note that spinlocks have always had a problem with priority inversion (OSSpinLock's back-off algorithm is just to call syscall_thread_switch with a 1ms delay every 1024 iterations, which is not bad but doesn't solve all priority inversions anyway).
The expression hurry up and wait comes to mind. I would rather have a thread sleeping while it waits for slightly less perceptual performance than a thread spinning and chewing through my battery because it's waiting as fast as it can.
(I had the same problem years ago with somebody's spinlocks on Windows running on a single-CPU system; the obvious functions to yield a timeslice will yield only to a thread of equal or higher priority, meaning you can starve a lower-priority thread. But even when this happens, the lower-priority thread does get a timeslice eventually - which sounds like the OS X situation.)
That obviously breaks the whole system, spinlocks or not (threads doing long CPU-bound computations will block everything else).
You can cause pretty bad performance by running too long in user-interactive or in OpenCL, so, like, don't do that.