The services I’m talking about are not as you describe. It’s usually very stateful, often involving a complex queue system that collects information about the user’s request and processes it in ways that alter that user’s internal representation for recommender and collaborative filtering systems. It’s not like an RPC to a pure function, but is a very large-scale and multi-service backend orchestrating between a big variety of different machine learning services.
> “Once you have 100,000 lines of code that all operate on that linked list, you are stuck with it.”
No, I think this really is not true, and as long as the new “performant” redesigned linked list that you want to swap in can offer the same API, then this is a relatively easy refactoring problem. I’ve actually worked on problems like this where some deeply embedded and pivotal piece of code needs to be refactored. This is a known entity kind of problem. Unpleasant, sure, but very straightforward.
You seem to discount the reverse version of this problem which I’ve found to be far more common and more nasty to refactor. In the reverse problem, instead of being “stuck” with a certain linked list, you end up being stuck with some mangled and indecipherable set of “critical sections” of code where someone does some unholy low-level performance optimization, and nobody is allowed to change it. You end up architecting huge chunks of the system around the required use of a particular data flow or a particular compiled extension module with a certain extra hacked data structure or something, and this stuff accumulates like kruft over time and gets deeply wedded into makefiles, macros, and deployment scripts, etc., all in the name of optimization.
Later on, when requirements change, or when the overall system faces new circumstances that might allow for trading away some performance in favor of a more optimized architecture or the use of a new third party tool or something, you just can’t, because the whole thing is a house of cards predicated on deep-seated assumptions about these kludged-together performance hacks. This is usually the death knell of that program and the point at which people reluctantly start looking into just rewriting it without allowing any overcommitment to optimization.
(One example where this was really terrible was some in-house optimizations for sparse matrices for text processing. Virtually every time someone wanted to extend those data structures and helper functions to match new features of comparable third-party tools, they would hit insane issues with unexpected performance problems, all boiling down to a chronic over-reliance on internal optimizations. It made the whole thing extremely inflexible. Finally, when the team decided to actually just route all our processing through the third party library anyway, it was then a huge chore to figure out what hacks could be stripped away and which ones were still needed. That particular lack of modularity truly was caused specifically by a suboptimal overcommitment to performance optimization.)
“Overcommitting” is usually a bad thing in software, whether it’s overcommitting to prioritize performance or overcommitting to prioritize convenience.
The difference is that trying to optimize performance ahead of time often leads to wasted effort, fast running things that don’t solve a problem anyone cares about anymore or lack that critical new feature which totally breaks the performance constraints.
Optimizing to make it easy to adapt, hack these things in, have a very extensible and malleable implementation is almost always the better bet because the business people will be happy you can accomodate rapid changes, and more able to negotiate compromises or delayed delivery dates if the flexibility runs into a performance problem that truly must be solved.