If it becomes the case that the hardware has changed in such a way that most programmers now have a hard time writing code that compiles efficiently to it, the hardware has changed in the wrong way! Well, that is the case over the past few decades! Hardware has changed in the wrong way.
What is wrong with shared memory? Also if HFT programmers can do it, wouldn't that imply that the hardware can do it?
Consider the simplest example of a busy-loop, which is likely close to the fastest you can get. Core1 sends a message to Core2 by writing to an address that Core2 is repeatedly reading from.
Core1 first has to gain exclusive access to the address, by broadcasting an invalidate on the bus. Core2 thus discards its cache line. Next Core1 writes to the address by modifying the data in its L1 cache (or store buffer or worse). Core2 then reads from the address. Since Core1 holds the modified cache line, it is required to snoop the read. Core1 tells Core2 to wait and then retry. This happens repeatedly until Core1's write lands in main memory. Now Core2 can read the memory.
So:
1. Messaging requires a round trip to memory, a bunch of cache line invalidation, and other nonsense. Messaging involves multiple transitions to and from shared state.
2. Your CPUs already have a very fast message bus to implement MESI, but it is unavailable to you as a client programmer.
It should be obvious from this that we could build a much faster CPU messaging architecture if we cared to.
Shared memory is extremely fast on modern x86.
Hardware message interfaces would go a good deal into improving the situation.
Currently, we just turn any random problem into a cache invalidation one, and hope it makes our life easier.