For example:
1) Why do we have entirely different management structures for anonymous memory (anon_vma and friends) and shared memory (inodes) when they really end up doing the same job? (Anonymous memory can be shared, so both paths end up needing to do the same kind of thing.) ISTM we can just use shared objects in all cases, tweaking the semantics as needed for shared memory. Anonymous memory is just swap-backed shared memory, after all. It should use the same logic.
(Yes, you need to work without a swapfile. No, that doesn't change the conceptual model.)
2) Do we really need open-coded page table walking? Why duplicate essentially the same logic four times? Sure, you occasionally want to do slightly different things at each page table level, but you can still unify most of the logic. (If you really want, you can encode the per-level differences in code generated via C++ template.)
3) Do micro-optimizations of the sort mentioned in the article really help? If these sequence numbers had been 64 bits long, they wouldn't have overflowed. If the VMA tree had been implemented as a splay tree (like on NT) instead of an RB tree with a per-thread cache, the per-thread cache might not have been needed. (Splay trees automatically cache recently-accessed nodes.) How much do these little tricks actually help? Is their cumulative effect positive?
4) Why is so much of the vm internal logic spread throughout the kernel? Both NT and the BSDs have a well-defined API that separates the MM subsystem from the rest of ring 0, but on Linux, ISTM that lots of code needs to care directly about low-level mm structures and locks. Why should code very far from the mm core (like the i915 driver) take mmap_sem directly? It feels like more abstraction would be possible here.
I get arguments based around ruthless pursuit of performance, but with function calls taking nanoseconds and page faults taking orders of magnitude more time, is anyone saving anything significant?