Hum... The hardware isn't actually there.
It has the capabilities the GP wants, but it doesn't export them in a way you could tune into that NUMA machine.
It has the capabilities the GP wants, but it doesn't export them in a way you could tune into that NUMA machine.
Sure you can't hide metadata in the unused bits of a pointer, but that didn't seem particularly critical for making good use of a NUMA machine.
Fast memory, addressable by the core.
The one that is there, but can only used by accessing the slow memory.
I think ultimately what I'm asking for means a much, much tighter bound on worst case performance, but at the cost of best case performance. That most of us never see anyway. For certain workloads, that could end up being a net positive. And there's probably some way to expose CPU metadata that gives some of that difference back.