- Data migrations, schema changes, backfills, backups & restores, etc., take so long that they can either cause or risk outages or just waste a ton of engineer time waiting around for operations to complete. If you have serious service level objectives regarding time to restore from backup, that alone could be a forcing function for horizontal sharding (doing a point-in-time backup of a 40TB database while dropping some unwanted bad DELETE transaction or something like that from the transaction log is going to be very slow and cause a long outage).
- The lack of fault isolation means that any rogue user or process making expensive queries impacts performance and availability for all users, vs being able to limit unavailability to a single shard
- When people don't have horizontal scalability, I've seen them normalize things like not using transactions and not using consistent reads even when both would substantially improve developer and end-user experience, with the explanation being a need to protect the primary/write database. It's kind of like being in an abusive relationship: you internalize the fear of overloading the primary/write server and start to think it's normal not to be able to consistently read back the data you just wrote at scale or not to be able to use transactions that span multiple tables or seconds as appropriate.
IME vertically scaled replicas/hot-stand-bys are a lot more stable to operate if your requirements allow you to get away with it. OtoH you better already be prepared if/when you hit scaling limits.
Vertical scaling is more expensive without the benefit of failover. With modern clouds, Horizontal scaling is cheaper with commodity hardware and gives the added benefit of failover redundancy.
So maybe one of the core reasons for Silicon Valley's love for horizontal scaling is due to plentiful AWS credits.
I can both horizontally and vertically scale in AWS. I can both horizontally and vertically scale in my home office.
Horizontally scaling is just more cost efficient if you can dynamically adjust to load. Clouds enable that because they have massive reserve compute.
In the startup world, there infinite compute, but finite money. It’s cost savings.
If you're using cloud stuff like EC2 on AWS for cost savings, you're in for a bad time. It makes sense for stuff that you need to be able to dynamically scale on a whim, but once you know what kind of setup you need, it's almost always cheaper to go dedicated.
At AWS's data center AWS also buys big servers. But their advantage is that they only have big servers, and then rent it out in pieces. They don't care whether they rent you 2 48 core 96GB instances or one 96 core 192GB instances. Both options run as VMs on much larger servers and take the same resources, so both are priced the same (a machine of the same category with twice the resources costs exactly twice as much). Thus there's a much smaller benefit to vertical scaling.
If you compare actual prices on large-ish instance types between AWS and regular dedicated servers (rented in a DC, not doing your own thing) then AWS is very expensive. And the hopes of paying less due to better utilization somehow never pan out. AWS has lots of advantages, but price isn't one of them; and their price structure will influence the architecture you choose.
That being said I'm partial of having fleets as small as possible to simplify operations. While trying different instance sizes, I found our bottleneck ended up being network. The largest instance types saturate the network first, and hence we can't use servers as efficiently.
Those big ass SQL boxes cost serious money.