And you don't get to waste time scrolling, your typical module's code fits on your single screen.
Anything about the L1 cache and K is just wrong, usually. At 600KB the K4 database engine is much too large to fit in L1 (K9 from Shakti is somewhat smaller but still a few times too large). And L1 instruction cache misses aren't a bottleneck for other languages, so there's little benefit in reducing them even to the extent K does it.
The long version: https://mlochbaum.github.io/BQN/implementation/kclaims.html
Is that really the bottleneck? I've done quite a lot of profiling on high performance code and I've almost never hit a bottleneck in the instruction cache. Data access bottlenecks or branching hit performance harder and sooner than instruction fetching.
> And you don't get to waste time scrolling, your typical module's code fits on your single screen.
How much of the time you save scrolling is spent on decoding an array of symbols and remembering what those symbols are?