It's probably because if your data is spanned across two pages, it may require a lot of page swaps if you access data back and forth?
We can do OoO execution, why can't we do OoO prefetching w.r.t. page faults? (I.e. try to fetch into a cache, if there would be a page fault don't. If there's something that would cause that fetch to page fault between it being prefetched and it logically being fetched, invalidate the cace.)
OoO engine were able to issue address calculation from loads (if address register is ready) ahead and thusly when real load instruction execute the data is already in cache or closer to cache.
It is really easy to do, actually. It can even simplify page fault handling in CPU.
That being said if you're really looking to optimize for TLB misses you should be using huge pages if your OS and processor support them.