You're more exact, I tried to simplify.
Whether they're CPU local depends on cache line state. It's fast in MESI protocol exclusive or modified state. Slow in invalid and shared states.
However real systems have contended objects pretty much always.
I'm vaguely aware about differential reference counting, but don't know any systems that use it. A bit like fast multicore counters, that break one counter into multiple uncontended "sub-counters". Read operation adds all sub-counters together to get the "real" value.
> Re alloc and free, that's completely orthogonal with refcounting.
Technically true, but in real systems those go hand in hand almost always. Yeah, I have written firmware for embedded devices that do refcounting and don't allocate memory. But that's not the usual case.