>If that is "hard", I would hate to see what you think of things like cache coherency across cores...
Cache coherence across cores is easy. The number of corner cases is much smaller.
The number of corner cases of old architectures (MIPS, or SPARC) is staggering. I designed a MIPS core prototype and I was done in single week with all arithmetic and memory access commands and spent three weeks designing, implementing and debugging branch handling, due to branch delay slot "feature".
>You do a check for the full range before the access, because for best performance you would want to pull in as much as possible in one go anyway.
You cannot do that - pull as much as possible in one go, - for case page boundary is crossed. But the very fact that you have to check before issuing operation make things more complex than they can be.
>>"load multiple" from PC (which is exposed to programmer in quite peculiar way)
>What's hard about that?
State handling. Can this instruction be interrupted?
>>Or out-of-order execution.
>Everything gets turned into uops, this instruction can emit more than one.
There are more than "everything gets turned into uops" out-of-order execution engines - scoreboarding, for example, provides most of benefits for much less energy cost.