Yeah, in my example, the $1k is for one server (cluster node). These servers (Purus, Quanta etc.) are commonly used to rapidly build enterprise clouds (you usually buy them by the rack). It's the closest thing you get to plugging a network cable into a bunch of Xeons. The cost of one system breaking is negligible. This is not counting the colo costs.
You can also do this with virtual public clouds (AWS, linode, GCP et al), but you'll of course pay a premium for the infrastructure. This might be worth it though, because you can now scale within seconds to handle qps bursts. Usually, latency can be lowered by going baremetal (see e.g. Algolia).
Academia should be able to handle more qps than our system, because the queries are really trivial in comparison. With decent caching, an 8c should be able to do 50 to 80qps. That's what I get from a few experiments when I switch my test cluster into restricted mode (basically just substring search).
Of course I can only speak from my experience, not how this can be applied to Academia's existing infrastructure. Testing large search engine deployments can be really, really frustrating.