Intel DDIO, LLC cache, buffer alignment, prefetching, and packet rates
adrianchadd.blogspot.com
adrianchadd.blogspot.com
There would be more circuitry to divide the addresses. I haven't done too much with hardware (only some light FPGA experience), so I don't know how much extra, but it doesn't feel like it would be prohibitive. I could be wrong.
(If you're aware of someone doing analysis on that, I'd be curious to see it -- I haven't seen any.)
You can accomplish division by a constant / modulus by a constant without a full division, however. For example, (x == x % 256 - x * 256), mod 257. That sort of thing.
http://www.cs.utexas.edu/~skeckler/pubs/MICRO_2014_Modulus_I...
They evaluate the technique on a GPU because bank and set conflicts can lead to bad performance more easily on a GPU than a CPU.
Interesting trick is using different values for cache index and tag (logical address for index and physical address as tag is common, but that's mostly about doing cache lookup concurrently with TLB lookup). In theory one could use result of some hash function as index, which would certainly change the aliasing behavior (although it is not exactly certain that such change would be for the better) and again, there is the additional cost of calculating said hash function (in combinatorial logic and with reasonably small delay).
http://lemire.me/blog/archives/2012/05/31/data-alignment-for...