On x86 multiprocessors, if, without taking any special measures, you set x=1 from processor 0, and processor 1 reads x shortly, but not too shortly, thereafter, it will see 1. If you try that on, say, an 80-core ARM CPU [1], CPU #53 might not see a 1 for a long time, if ever.[2] You're no longer guaranteed that cache changes propagate unasked. As we get more and more CPUs, the cost of creating the illusion that there's a single memory goes up. It's now important that compilers know what's shared so they can generate the proper fence and barrier instructions. Especially since the latest iteration of ARM has new features such as "non-temporal load and store."
What used to be a theoretical problem, or at most a problem for OS architects, has thus acquired teeth and claws that can bite ordinary applications programmers.
This is where traditional C++ mutexes start to break down. From a theoretical perspective, it's always been annoying that mutexes at the POSIX level have no tie to what they're supposed to be protecting. Hitting a mutex has to mean "flush everything". That's inefficient. Rust has the advantage that its locks own data, so the compiler knows what has to be flushed. That becomes more important as the number of CPU cores becomes very large.
I've said for years that the three issues in C/C++ are "how big is it", "who owns it", and "who locks it". C++ now at least has abstractions for all of these, although they all leak. You can still get raw pointers out to misuse, or, more likely, pass to some API for misuse.
[1] https://venturebeat.com/2020/03/03/ampere-altra-is-the-first...
[2] https://developer.arm.com/documentation/den0024/a/the-a64-in...