> If you imagine how a physical CPU or GPU has to be constructed in order to do large multi-input operations (...) You can imagine these inputs as being in "lanes" that are arranged across the chip such that the inputs to each lane are stored near the lane.
This is not how a GPU register bank works at all. GPU register file are SRAM banks and operand collector are used to handle register-read latency. And there is a big cross-bar between the register banks and the operands collectors.
And for CPU Vector unit and SIMD unit, I only know two implementations (the CVA6 ARA vector unit and an industrial closed source one) but neither of them do registers storage within/near the lane.
Author's assumption on microarchitecture seems questionable to me.
PS: The ARA RISC-V V implementation used a mask unit to handle mask. Which makes the mentioned problem irrelevant