The hardware that these programs are running on store objects in linear memory, so it doesn't not make sense to treat it as such.
Caches aren't quite as mix-and-match, but they can still internally manage different temporal versions of a cache line, as well as (hopefully) mask the fact that a write to DRAM from one core isn't an atomic operation instantly visible to all other cores.
Practice is always more complicated than theory.
It's still logically the same thing with these optimizations, obviously -- since they aren't supposed to change the logic.