i don’t believe rseq based cpu local caches require memory barriers on the fast path.
i don’t believe rseq based cpu local caches require memory barriers on the fast path.
rseq_cs
The rseq_cs field is a pointer to a struct rseq_cs. Is is NULL when no
rseq assembly block critical section is active for the registered
thread. Setting it to point to a critical section descriptor (struct
rseq_cs) marks the beginning of the critical section.
I’m not sure I fully understand that man page (it never seems to say callers have to clear that field at the end of a critical section, for example), but doesn’t that mean the caller has to guarantee setting rseq_cs happens_before any code in the critical section? That’s a memory barrier.No explicit memory barrier is required anywhere as the value is only read in supervisor mode and a privilege switch implicitly issues a LS-LS barrier on all major architectures. Even if you did not want to rely on that, you would only need a single S-LS barrier when you store the control structure the very first time.
Aha! So, to take advantage of that, a memory allocator uses the same abort handler for all operations?
If `rseq_cs` is no longer describing a relevant address, that is, the program counter has moved past it, the kernel just ignores it.
The code is executed by a single thread. Everything retires in program order. There is no need for a memory barrier between starting the critical section and its body.
Meanwhile the better the scheduler performs the more competitive the thread local approach becomes.
cost for interruption in an rseq critical section is that the PC gets overwritten to the rseq abort entry point before the task is rescheduled. no management thread necessary.
should be fairly minimal cost, especially assuming interruptions in the critical section are rare.
Giving it some more thought, I expect caches will typically be wiped out by a context switch. So the only place rseq is likely to benefit an allocator is on systems with multiple NUMA nodes where you'd like to make sure any management code isn't paying a penalty by hitting the wrong address range.
IIUC rseq (ie CPU local data) is primarily good for two things. The first being obviating the need for atomics (specifically the resultant cache line ping-pong) but thread local data already accomplishes that. The second being massive oversubscription of physical CPU cores (ie tens of thousands of threads) where TLS becomes utterly wasteful while also thrashing the cache.
maybe i give it too much weight, but:
> massive oversubscription of physical CPU cores (ie tens of thousands of threads) where TLS becomes utterly wasteful while also thrashing the cache
seems worth solving to me
If we do decide to concern ourselves with performance TLS has zero overhead and doesn't suffer from contention while rseq (at minimum) carries a penalty if preempted and involves setting a flag plus exhibits a data dependence for the address offset (the latter since AFAIK compilers don't natively support it as they do TLS). So while I'm certainly open to benchmarks to me it very much looks like a mixed bag that only comes up when you're already in questionable territory to begin with. In the event that we do step outside the norm I'd guess that a handful of threads with contention is a much more common scenario than thousands of threads per physical core exhibiting only minimal preemption.
I also expect something like a web server servicing thousands of requests in parallel to use an event loop instead of spawning an equivalent number of threads. I'm struggling to come up with a scenario where you haven't already fatally shot yourself in the foot and this remains a useful optimization to make. It's certainly relevant if you're using fibers (green threads, whatever you want to call them) but at that point you aren't in c calling malloc and your language runtime will (one hopes) already be taking care of all this for you.
Totally fair, and I agree that you should only have a few threads (or TPC) and shouldn’t be inundating the allocator with requests. But I’m only grinding this axe because I’ve personally had to deal with a system that went against most of that guidance.
We ran far too many threads in a memory-constrained environment. Thread count was many multiples of core count.
I totally agree that this is “questionable territory”, but honestly, any application that’s outgrown the basic glibc malloc has made a few mistakes.