Shipping a monolith worked on by 250+ devs in 30 teams is slow due to the coordination needed as compared to 30 teams shipping 50-75 services.
Vertical scaling can take you really, really far operationally.
But again, most systems will never be there ;)
Coming up with a greenfield microservice design with arbitrary responsibilities and intercommunication feels so stupid to me. Why not build the thing as a monolith and split parts out when you actually have scaling problems, instead of solving theoretical problems? Development velocity is going to be way higher and it gives developers a chance to discover problems without having to deal with the mental overhead of shipping N services all at once.
The value of fast iteration cannot be overstated. Build the minimum viable version of the project first, then stress test it and break parts out when needed. This is much easier if you write modular code.
I sometimes think programmers are the last people who should be writing software.
The personality type that likes writing code is the exact type that likes tinkering around the edge and working on hypotheticals instead of addressing the problem at hand.
Good programmers pride themselves on striking compromises and shipping a smaller thing sooner and they love iterating. The ones that are working on hypotheticals unchecked are not bad--just nobody has educated them.
Then, uncoincidentally, they started crying about how they couldn't find work anymore.
- "I'm struggling to find work and need to make money. What can I do?"
- "Learn to code, good buddy! Software engineering solves all problems."
2025:
- "I learned to code and am still struggling to find work and need to make money."
- "You shouldn't be doing software engineering."
This happens at all sorts of companies that are not FAANGs, and don't have "insane scale". There's an extremely large spectrum between the web site for Joe's Coffee Bar and Amazon. On commodity hardware, you hit issues with scaling up a monolith long before you reach FAANG level or "insane scale".
I've worked with multiple startups that have hit scaling limits with their monoliths. Inevitably, dealing with that is a huge problem because the monolith was developed with few resources under heavy time pressure. Modularity is lacking, breaking it up is difficult. Individual devs are often inclined to say that's just a skill issue, and that may have some truth to it, but managing those skill issues is a big part of what corporate software development is about.
This can have a huge impact on a company's funding, ability to deliver new features, ability to scale development, and of course ability to scale the user base. Typically, by the time they hit that wall, scaling the monolith horizontally is not a great option, because it wasn't designed to support that.
It's often been observed that microservices are a primarily an organizational tool, and that's true. But organization is critical if you have multiple development teams.
That doesn't necessarily mean every app should consist of hundreds of tiny microservices. But there can be enormous benefits from implementing an app from the start as independent services based on its natural divisions between modules.
The tooling is good enough to scale out, but micro services are mostly beneficial for organizational scaling.
The value for other concerns is mostly situational.
Unless you have some very specific need, I was horizontally scaling the app layer in 1996 even in true monoliths.
K8s can be too much, but even when tooling and costs forced us to segment by technology layers, you never tried two node failovers.
If you are trying to share state at the app level you would most reduce availability, because of split brain etc…
The persistence layer was bad enough with shared quorum drives, heartbeat networks etc…
It sounds like you are just spinning up to app servers or are you talking about active/passive or active/active two node clusters?
That is vertically scaling in the way I understand the term, not just scaling out two instances.
If you're message driven you can maybe have a set of replicas consume off of a message bus like Kafka or NATS, but now you need to maintain that cluster, as well as build a messaging layer. What protocol do you use, etc?
Then there's the question of what triggers scaling up/down. The simple way is scale on CPU utilisation but then you get flapping issues. Or if load isn't completely evenly distributed you get one replica starving for CPU whilst the rest have too much.
All of these questions go away if you can just provision a big node and be comfortable that it will be able to handle anything within reason.