1. Sometimes business needs require non-linear growth quickly. One day your perfectly optimized process now requires an n-squared algo across billions of records. Suddenly your one machine is tiny compared to what it used to be.
2. If you haven't scaled the workload horizontally, it can lead to a rewrite just to begin to distribute the load. When you hit the limits of a single machine, it is often a hard barrier that you cannot easily cross.
3. Distributed systems are distributing 3 main resources: memory, IO, and computation. IO especially quickly becomes a bottleneck on a single system, but memory is also quite thorny because memory management itself can be a bottleneck on a vertically scaled system.
4. People with distributed systems do obsess over single-node QPS. If it takes 5k nodes to do work and you can optimize down to 2k nodes, you are saving a lot of money! However, it isn't that simple and this is where being properly distributed gives you cost leverage. You might find that 5 highly scaled machines are more costly than 50 commodity machines, especially in the cloud ecosystems.
5. Finally, and probably to your point and the GP's points, you kind of have to go with the flow when it comes to the level of abstraction people are writing the code at. It takes increasingly specialized knowledge to optimize a process that will run well on a 100gb process (e.g. virtual machine garbage collection issues). You are knowingly sacrificing efficiency for the nice higher level abstractions.
That said, I would emphasize that the abstractions become a smaller slice of the performance pie when you're dealing with algorithmic complexity. C won't magically make your Python algo O(1).
I don't disagree with your main point btw. I think a well done processing / data pipeline should have distributed systems available and single-node computational scenarios available, because there is an undeniable complexity gain when you reach for a distributed system immediately.
I would argue, however, that it has little to do with the decision of whether or not to distribute your system. I've spent a lot of time dealing, for example, with bottlenecked single instance RDBMS instances that are handling load they shouldn't be handling. (For example those accidental recursive queries that are often a side effect of nice ORM abstractions) I totally agree with you, and have seen it happen, that people who do not understand the performance characteristics of their system can reach for a distributed solution before they've understood what their performance issue was. But I've also seen plenty of situations where they reach for bigger hardware for the same reasons.
I think deciding whether or not to distribute means taking a disciplined approach to projecting the business needs of particular data entities you'll be dealing with. For example, if you have a users table, and it will ultimately store every human in the United States, then, well, that is quite do-able on today's single instance RDBMS systems, and you can project the theoretical growth over time. And if you need read load and HA, then you can go a replication route, or at least look at that first before doing something that reduces the quality of your transaction handling, like sharding. And then once the system is in place, taking a disciplined approach to profiling and quantifying the costs of the different aspects of the system and justifying their business value. For example having run large scale recommender systems, a typical decision might be to degrade the quality of an algorithm if it means saving a tremendous amount of money on processing.
What I'm talking about is like 3,000 nodes to serve 15,000 QPS. That seems like a lot of QPS, and so it makes intuitive sense to people that they'd need a big distributed system for it, unless they have a sense of how fast computers actually are.
A classic example of a "cheap" service is something like Stack Overflow [1] or Wikipedia: serving content that is mostly static; you just gotta configure your caches well. A classic example of an "expensive" service: a search engine, where so many queries cannot be cached (too rare), or are even completely unique (never seen before [2]).
[1] https://stackexchange.com/performance
[2] https://www.seroundtable.com/google-15-percent-queries-25730...
[1] https://twitter.com/Nick_Craver/status/1280494336673751044?s...
The places I’ve worked (small/med companies) with lots of instances were always well utilized. People understood well the cost of infrastructure and put effort into minimizing it. But as features and services and customers grow, so do instances.