Huh, this is an interesting instruction. Wonder if this will “break” code that expected a read into an unmapped page to fault…
Huh, this is an interesting instruction. Wonder if this will “break” code that expected a read into an unmapped page to fault…
If you perform a page-crossing load where, say, only elements 0-5 out of 32 elements are loaded, then only elements 0-5 in the predicate register gets set, the loop steps 6 elements instead of 32 elements in this iteration. On the next time around, if you really intended to go past the end, byte 0 of the next page will be in the first element position and goes through normal fault handling.
Lemire's example loop is subtly wrong here, but the version in the SVE slides he links to is fine.
According to some documentation[0], anything after the first element uses the `MemNF` "intrinsic":
if first then
// Mem[] will not return if a fault is detected for the first active element
data = Mem[addr, mbytes, AccType_NORMAL];
first = FALSE;
else
// MemNF[] will return fault=TRUE if access is not performed for any reason
(data, fault) = MemNF[addr, mbytes, AccType_CNOTFIRST];
However, if the data would fault, said "intrinsic"[1] returns `UNKNOWN` for the data: (value<7:0>, bad) = MemSingleNF[address, 1, acctype, aligned];
if bad then
return (bits(8*size) UNKNOWN, TRUE);
But clearly, the value of `UNKNOWN` is visible to the program, so what is it? I'd assume it would either be zeros or whatever the previous scalar was (i.e. untouched), but does anyone know?[0]: http://hehezhou.cn/isa/ldff1b_z_p_br.html
[1]: http://hehezhou.cn/isa/shared_pseudocode.html#impl-aarch64.M...
It is a fair assumption these days that no instruction leaves bits in the register "untouched" unless doing so is a core part of what's needed for the instruction to work. On architectures with register renaming (which is all high-performance architectures) it can be surprisingly expensive.
The very least, it forces a dependency on all previous instructions that write that register, and adds another source operand to the instruction. And on most real architectures, this source operand would be required to be present before the instruction itself can start. This is particularly bad for load instructions, if it pushes back when the load can be submitted to cache. This can wipe out all your memory parallelism, and effectively cut your performance to a fraction of what it would othewise be.
The basic mental model you need to have is that every instruction that writes to a register always gets a new register that starts with all bits zeroed. To have anything else, you have to do work.
- Whether subsequent non-faulting elements are loaded as normal, or ignored and treated as also faulting (more relevant for the FF gather variants)
- Whether the faulting elements zero or merge the previous element value of the destination vector
The first one has an awesome typo (I hope, else they're being a bit evil) with a double negative in the text explaining the instruction:
Inactive elements will not not cause a read from Device memory or signal a fault, and are set to zero in the destination vector.
Note the double "not" in the first half. Ouch.
Certainly, “contents of unmapped region” wouldn’t be on the table these days, right?