Could this conceivably be used to speed binaries like nginx? I tend to compile nginx on my systems as over time there are always custom modules I want to add/modify anyway.
Also low cache levels have less latency and yet are much smaller, the L1 instruction cache is 32KB. Any linear access of memory will prefetch and minimize the latency of memory access.