I came by to say essentially the same thing. The article is complaining that CPU behavior is treated as axiomatic by software but that treatment is flawed because the CPU may not satisfy all the conditions of its assumed behavior.
I have had the experience of seeing the 'sausage making' of CPUs on the other side when I worked at Intel and the verification of the 80286 was a big thing and the evolution of the 80386 was just starting. And the sad truth is that hardware engineers, exactly like software engineers, are often compelled to accept a 'good enough' solution when the business decides they have reached that point of what the business is willing to invest against the expected monetary returns.
In software this will express itself as an algorithm that is correct 99.99% of the time. Or one which is computationally expensive and so replaced with an algorithm that takes less computation but gives good enough results. And in the normal course of things, those decisions are often "ok" because failures are rare and when they fail there isn't any "gain" by them failing.
Security of course turns that calculus on its head, where, if you can break the chain somewhere, you win. So a place where the chain is only 99.9% as strong as it is elsewhere will get special attention. And when the chain breaks there is gain that exceeds the effort involved.
For much of the behavior that is responsible for both Spectre and Meltdown, the CPU can be constructed to not bypass the other secure checks. If the designer chose to, they could have the CPU machine check or fault if it read protected memory in a speculative path. Even if that path wasn't actually taken.
I don't know of course, but I would expect the typical conversation at a CPU design house would be, this;
A) We should always fault if we access protected memory
B) But what if its just an uninitialized pointer on a branch that won't be taken?
A) It is still a fault.
B) Ok but really "only" if we take the branch so lets wait until the code is actually executed to fault. Tree in the woods with no one around to hear it and all that ...
Good engineers, designing cool machines, making completely reasonable decisions about what is 'good enough', against a spec that says "CPU will fault if it ever accesses protected memory." because it will fault if the code ever gets there. As opposed to designing the fetcher to always check memory protections and always fault. To any reasonable person, both would seem to be sufficient, but only one is correct. And since it takes more silicon and more complexity to always check, we get the one that is good enough. It is my hope that at this very moment there are CPU designers all around the world changing the HDL in their design to be correct, and verifying that they run just as fast as they did before, just that now your CPU can fault on an address that wasn't actually executed. I am looking for any example where speculative execution or fetching of protected memory is ever not a 'bug' waiting to happen in a piece of code.
As a result I expect CPU designs to be updated, and new micro-architecture where this behavior always faults but is otherwise just as performant. No doubt it will take some additional transistors.