To do it correctly, lock needs to be done in the kernel thus obtaining a lock requires calling into the kernel which is more expensive.
I think you meant the memory barrier for syncing cache is just as expensive as the lock version, which is true.
Otherwise you inspect the implementation, but in 2024 a fast-pathed OS lock is table stakes.
The Windows and Linux solutions are by Mara Bos (the MacOS one might be too, I don't know)
The Windows one is very elegant but opaque. Basically Microsoft provides an appropriate API ("Slim Reader/Writer Locks") and Mara's code just uses that API.
The Linux one shows exactly how to use a Futex: if you know what a futex is, yeah, Rust just uses a futex. If you don't, go read about the Futex, it's clever.