When that trick can't be used, I think the most efficient method would be to clamp the top of the address so that the max would land on a single guard page. On x86-64 and ARM that would be done with a `cmp`and a conditional move. RISC-V (RVA22 profile and up) has a `max` instruction. That would be typically one additional cycle.
The new proposal is for using a 64-bit pointer and a 64-bit offset, which would create a 65-bit effective index. So neither method above could be used. I think each address calculation would first have to add and check for overflow, then do the bounds-check, and then add the base address to the linear memory.