I'm talking about situations like search, where you hold the entire index in ram. Total # of machines = size of index / indexserver ram. Usually the apps that run on these have, say, 96 cores and they're using about 80, and the idle time is mostly instructions waiting for memory fetches.
Typically that index fronts a disk repository which wouldn't fit in RAM, although over time, what fit in RAM, what fit in disk, what lived in RAM, what got cached in RAM, etc, have changed over time.
BTW, I'm probably of the same generation as you and the single most important lesson I ever learned for computing performance was "add more RAM"; in the days when I first started using Linux with 4MB of RAM, it wasn't enough to do X11, g++ and emacs all at the same time without swapping, so I spent my hard-earned money to max out the RAM, at which point it didn't swap and I could actually do software development quickly.