I’d be interested to hear more about this.
I’d be interested to hear more about this.
He was comparing to two CPUs sharing data with shared memory. In the shared memory case, each CPU stays on the happy path but sometimes has to do some cache coherence, which may stall the CPU, but that’s the worst that can happen, generally.
But if you try to send a message to another CPU, then you’re asking it to raise an interrupt. That’s not the happy path. CPU will have to stop everything to receive the interrupt and then divert execution to some interrupt gate. Arvind’s point was that every impl of this is going to be much much worse than the worst case of cache coherence.
Every measurement I’ve ever made confirms this.
https://par.nsf.gov/servlets/purl/10079614
To my knowledge, the remaining cost could be decreased to approximately the same cost as a branch mispredict, but getting there would require changes to the chip hardware and software stack.
Do it even need to be a misprediction?
If you are completely focused on latency then flushing everything else makes sense. But I would think that if you continue execution for now and put a branch instruction into the queue you'd reduce the cost per interrupt even further.